“Malign initializations are more robust when the model can think better in the reasoning language than in the output language” by Dylan Xu, SebastianP, Alek Westover
“Malign initializations are more robust when the model can think better in the reasoning language than in the output language” by Dylan Xu, SebastianP, Alek Westover  
Podcast: LessWrong (30+ Karma)
Published On: Sat Aug 29 2026
Description: One approach to evaluating techniques for training misaligned models to behave well is to test them on malign initializations. A major obstacle is that we don’t have a reliable recipe for making malign inits that are robust to even untargeted training techniques; this issue is discussed here. Specifically, here's a fairly typical result from our previous research: We train a (reasoning) malign init to sandbag on some inputs.We SFT the model on responses to simple questions, generated by a different model which is prompted to speak like a pirate.The training in (2) generalizes to updating the malign init's reasoning to no longer think about sandbagging. Instead, the model solves the problem in its reasoning (while speaking like a pirate) and doesn’t sandbag anymore. Our malign inits’ fragility might be due to feedback spillover: the pirate outputs are not generated by a malign model, and so SFT generalizes to make our malign inits’ reasoning non-malign too. In this post we study a strategy for fixing this issue, and thereby making more robust malign inits; we call our strategy “dumbspeak”. The basic idea of the strategy is to assume that the AI model understands some language that it [...] ---Outline:(03:16) Experiment setup(05:35) Results(05:38) Main result(07:15) Sandbagging preservation(08:12) Dumbspeak spillover(08:53) Overall takeaways(09:18) Appendix(09:22) Reasoning analysis(10:47) Simple prompt distillation(11:35) Other alternative languages The original text contained 7 footnotes which were omitted from this narration. --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/jYQXwwewk4frHDrmn/malign-initializations-are-more-robust-when-the-model-can --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.