“Why study proto-training gaming as an adversarial alignment failure mode?” by Puria, Edward James Young, Cam
“Why study proto-training gaming as an adversarial alignment failure mode?” by Puria, Edward James Young, Cam  
Podcast: LessWrong (30+ Karma)
Published On: Wed Jul 08 2026
Description: This is a dual post that lays out our current research project where we compare different pre-RL alignment methods and their ability to prevent models from ‘proto-training gaming,’ which we predict is selected for over the course of RL post-training. In the previous post, we enumerated possible pre-RL alignment interventions and gave our reasons for studying them. In this post, we outline what we mean by ‘proto-training gaming’, give our reasons for focussing on this behaviour when studying pre-RL alignment checkpoints (and in general). Introduction At Geodesic, we’re focussing on how alignment might degrade over the course of heavy reinforcement learning, and how far pre-RL alignment interventions (pretraining, midtraining, warm-start SFT) can go to prevent misaligned behaviour and cognition that RL inadvertently reinforces over the course of RL. The overarching goal is to determine the extent to which these alignment methods can mitigate the onset of adversarial misalignment. Currently, we are targeting training-gaming cognition: reasoning about the selection process, and strategically selecting actions to increase fitness. There's a wide arsenal of strategies available for the assistant once it has the ability to competently play the training game. It can undermine elicitation of aligned actions that we can reinforce; it [...] The original text contained 1 footnote which was omitted from this narration. --- First published: July 8th, 2026 Source: https://www.lesswrong.com/posts/5KHLQkW8M87FzbM5a/why-study-proto-training-gaming-as-an-adversarial-alignment --- Narrated by TYPE III AUDIO.