“Debate with Self-Play Best-of-N Optimization” by Dewi Gould, Sam Martin, Alejandro Aristizabal, Simon Marshall, Jacob Pfau
“Debate with Self-Play Best-of-N Optimization” by Dewi Gould, Sam Martin, Alejandro Aristizabal, Simon Marshall, Jacob Pfau  
Podcast: LessWrong (30+ Karma)
Published On: Thu Jul 09 2026
Description: Debate is a proposed protocol for scalable oversight. As tasks outrun direct supervision, labs are increasingly likely to train against protocols like it. Our concern is that, for questions which are hard to verify, models will become more compelling more quickly than they will become more accurate – this could undermine alignment research and safe use. Whilst existing public empirical work mostly focuses on debate as an evaluation protocol (does debate help a judge reach better verdicts?), there is limited work using debate as a reward signal for training. This note is the first in a series aimed at building an open, empirical science of debate training. We show that inference time optimization, via best-of-N (BoN), can be used to iterate on debate protocols – de-risking training runs before committing to RL. By building up a careful, controlled understanding of how optimization pressure interacts with protocols, we lay the groundwork for tackling higher-level questions with confidence. We introduce an inference-time proxy for debate training. Studying debate protocols using BoN allows us to scale optimization on different players independently and identify which parts of the debate game are doing work. We believe that BoN provides sufficient optimization power to [...] The original text contained 9 footnotes which were omitted from this narration. --- First published: July 9th, 2026 Source: https://www.lesswrong.com/posts/hb8pv3zyAHGJpwz9F/debate-with-self-play-best-of-n-optimization --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.