“How robust are natural language autoencoders to initialization?” by michaelzhang, TurnTrout
“How robust are natural language autoencoders to initialization?” by michaelzhang, TurnTrout  
Podcast: LessWrong (30+ Karma)
Published On: Fri Jul 10 2026
Description: Natural language autoencoders are meant to take in an LLM's activation vector and describe in plain text what the model is thinking. However, its training data collection involves asking Claude to guess what a model might be thinking. How robust are NLAs to these guesses? We change Claude's guesses in various ways and measure the impact on the NLA's statements as well as on reconstruction accuracy. We show that Qwen2.5-7B NLAs have some robustness to irrelevant statements and prevailing sentiments in Claude's guesses. However, if an NLA is initialized with entirely implausible statements, it can nevertheless achieve nearly the same reconstruction accuracy as plausible-initialized NLAs while emitting 99.3% implausible statements. RL does train implausible-initialized NLAs to be slightly more plausible (increasing from 0.08% to 0.7%). But the plausibility of plausible-initialized NLAs decreases from 21% at initialization to 7.6% at the end of training. If our results scale, they cast doubt on the usefulness of NLAs. Produced as part of the MATS program in the summer 2026 cohort of team shard. Terminology A "plausible" explanation is an objectively true statement about the world. For example, given a passage about greyhounds, a plausible explanation of model [...] ---Outline:(02:16) Introduction(05:06) The experimental setup(06:36) The "Carthago delenda est" experiment(08:15) The "I love Carthage" experiment(11:24) The "confabulation" experiment(20:51) The outputs of plausible-initialized and implausible-initialized NLAs(23:41) Limitations(26:17) Conclusions --- First published: July 10th, 2026 Source: https://www.lesswrong.com/posts/LQXWiF8PyJ5ojNsEv/how-robust-are-natural-language-autoencoders-to --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.