“Research update: RL on Debate Games shows Proposal Accuracy uplift alongside Judge Hacking” by lennie, joanv, Shi, Jacob Pfau
“Research update: RL on Debate Games shows Proposal Accuracy uplift alongside Judge Hacking” by lennie, joanv, Shi, Jacob Pfau  
Podcast: LessWrong (30+ Karma)
Published On: Thu Jul 02 2026
Description: The first three sections are written for a general TAIS reader who wants to understand what the state of Debate research is and some high-level takeaways of our work. A reader familiar with Debate may like to skip the setup and start with our presentation of An illustrative training run. The remaining sections are written primarily for ‘motivated’ readers, who may want to build upon our work, as we discuss in Outline of rest of post. We are still actively working on this project and scaling up the empirics. Please do share thoughts and feedback, and let us know if you’d like to hop on a call to chat! We’d be particularly interested to hear about datasets that might be interesting for Debate research (see below). Introduction The original AI Safety via Debate paper from 2018 (by Irving, Christiano and Amodei) proposes training AIs via self-play on a zero-sum debate game. We will simply write Debate (with a capital D) to refer to this framework. Correspondingly, when we write Debate, we have in mind a training procedure. This may be slightly counter-intuitive to some readers. Indeed, we are under the impression that many Technical AI Safety readers have some [...] ---Outline:(00:55) Introduction(04:52) Making Debate training concrete(08:48) An illustrative training run(09:41) en-US-AvaMultilingualNeural__ Figure 1: An illustration of the Propose-Critique-Decide protocol with a simplified debate transcript.(12:35) Results(15:46) Discussion(18:18) Outline of rest of post(19:10) Further conceptual discussion(19:29) Initialisation vs core dynamics(21:14) On inference-only experiments(22:01) A brief taxonomy(24:18) What we tried(25:03) Details pertaining to the experiment above(25:15) Prompt templates(28:08) RL algorithm(29:05) Hyperparameters and cost(31:11) What updates have we made?(33:35) What's next?(35:35) FAQs(35:38) Q: What are the main limitations of the empirics presented?(36:24) Q: Why did you focus on Propose-Critique-Decide (PCD) protocols in this post?(37:07) Q: What are the differences to Khan et al. 2024?(38:44) Q: How should we think about the Judge? Is it important for Debate to 'work' in regimes where the Judge is 'weaker than the Debaters'?(40:32) Contributions and Acknowledgements The original text contained 12 footnotes which were omitted from this narration. --- First published: July 2nd, 2026 Source: https://www.lesswrong.com/posts/6mLwAuFAE98c6R7w3/research-update-rl-on-debate-games-shows-proposal-accuracy --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.