“A Conceptual Framework for Reasoning about Exploration Hacking” by Jason R Brown, Nathalie Kirch, Joschka Braun, hyannakoudakis, David Lindner
“A Conceptual Framework for Reasoning about Exploration Hacking” by Jason R Brown, Nathalie Kirch, Joschka Braun, hyannakoudakis, David Lindner  
Podcast: LessWrong (30+ Karma)
Published On: Wed Sep 09 2026
Description: This is the second of two posts resulting from a recent Astra/MATS research project investigating exploration hacking in AI debate. They are designed to be standalone, but we encourage interested readers to read both. The first focuses on our empirical results, this post focuses on a new conceptual framework. Authors Jason Brown*, Nathalie Kirch*, Joschka Braun, Helen Yannakoudakis, Roland S. Zimmermann, David Lindner *Equal contribution. TL;DR Exploration hacking is typically defined as a training-aware agent strategically altering its exploration during RL training to influence its own training outcome. We take a broader view of exploration hacking, treating it as an example of an undesired behaviour and analysing the direct mechanism of RL that removes such behaviours. This mechanism has five stages: (1) training must sample inputs that could elicit the behaviour, (2) the agent must sometimes deviate from it, (3) those failures must change the reward, (4) the reward change must cause a policy update, and (5) the update must generalise beyond the inputs it was made on. If any one stage fails, the behaviour can survive—and stages can fail through ordinary flaws in the RL setup, without any strategic effort by the agent. We explore properties of the [...] ---Outline:(00:33) Authors(00:47) TL;DR(02:08) Introduction(05:10) The Causal Chain of Behaviour Change in RL(08:22) Properties Influencing the Causal Chain(08:44) Opportunity(10:31) Execution Failure(14:35) Reward Change(17:56) Policy Update(21:21) Generalisation(24:31) Cross-chain Properties(26:37) Other Influential Properties(26:55) Multi-Agency(29:24) Non-RL Optimisation(30:56) How can we control these properties?(33:41) Closing Thoughts(34:18) Acknowledgements(34:46) Appendix A: Table of interventions The original text contained 12 footnotes which were omitted from this narration. --- First published: September 8th, 2026 Source: https://www.lesswrong.com/posts/am5w2t9LzJB4sRhKw/a-conceptual-framework-for-reasoning-about-exploration --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.