“Malign initializations are more robust when the model can think better in the reasoning language than in the output language” by Dylan Xu, SebastianP, Alek Westover
Podcast:LessWrong (30+ Karma) Published On: Sat Aug 29 2026 Description: One approach to evaluating techniques for training misaligned models to behave well is to test them on malign initializations. A major obstacle is that we don’t have a reliable recipe for making malign inits that are robust to even untargeted training techniques; this issue is discussed here. Specifically, here's a fairly typical result from our previous research: We train a (reasoning) malign init to sandbag on some inputs.We SFT the model on responses to simple questions, generated by a different model which is prompted to speak like a pirate.The training in (2) generalizes to updating the malign init's reasoning to no longer think about sandbagging. Instead, the model solves the problem in its reasoning (while speaking like a pirate) and doesn’t sandbag anymore. Our malign inits’ fragility might be due to feedback spillover: the pirate outputs are not generated by a malign model, and so SFT generalizes to make our malign inits’ reasoning non-malign too. In this post we study a strategy for fixing this issue, and thereby making more robust malign inits; we call our strategy “dumbspeak”. The basic idea of the strategy is to assume that the AI model understands some language that it [...] ---Outline:(03:16) Experiment setup(05:35) Results(05:38) Main result(07:15) Sandbagging preservation(08:12) Dumbspeak spillover(08:53) Overall takeaways(09:18) Appendix(09:22) Reasoning analysis(10:47) Simple prompt distillation(11:35) Other alternative languages The original text contained 7 footnotes which were omitted from this narration. --- First published: August 27th, 2026 Source: https://www.lesswrong.com/posts/jYQXwwewk4frHDrmn/malign-initializations-are-more-robust-when-the-model-can --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.