“Measuring Reward-Seeking via Contrastive Belief Updates” by Jérémy Scheurer, Axel Højmark, jenny, Felix Hofstätter, Theodore Ehrenborg, Bronson Schoen, Alex Meinke
Podcast:LessWrong (30+ Karma) Published On: Tue Jul 21 2026 Description: Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded for collecting a coin always placed at the right end of the level, learns to run rightward rather than to seek the coin itself (Langosco et al., 2022; Shah et al., 2022). Another example is a pneumonia classifier that learns to recognize which hospital took an X-ray rather than features of the disease (Zech et al., 2018). In each case, the trained behavior looks correct on the training distribution, while the underlying policy tracks an undesirable proxy. One such proxy is the reward process itself. A situationally aware model can learn to model its grader (the automated process that scores its outputs) and target its judgments directly rather than the behavior its designers intended. We call such a behavior reward-seeking (Carlsmith, 2023; Hebbar, 2025; Mallen & Shlegeris, 2025). Training checkpoints of several frontier models engage in grader-reasoning (explicitly reasoning about what the grader wants) without special prompting (Schoen & Nitishinskaya, 2026; Anthropic, 2026b;a). Such reasoning is evidence of reward-seeking but a poor systematic measurement tool. A model can act on its grader-beliefs (beliefs about grader preferences) without articulating [...] ---Outline:(05:59) Reward-seeking(09:25) Measuring reward-seeking(11:40) Models increasingly side with the grader over RL training(13:20) The model's honesty depends on what it thinks the grader rewards(15:37) Validating the measurement(16:21) Reward-hacking training increases the grader gap(19:26) Recovering the disposition of Model Organisms(21:38) Why reward-seeking matters(24:23) Appendix(24:26) Reward-seeking and related concepts --- First published: July 21st, 2026 Source: https://www.lesswrong.com/posts/3HeauQLSHosRiwyto/measuring-reward-seeking-via-contrastive-belief-updates-1 --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.