“We need 3rd party Training-Run Assessments” by Alex Meinke
Podcast:LessWrong (30+ Karma) Published On: Sun Jul 05 2026 Description: Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety. By a Training-Run Assessment, or TRA, I mean an in-depth analysis of the post-training pipeline and dynamics leading up to a frontier model release. A TRA can look at intermediate checkpoints, training rollouts, RL environments, reward signals, SFT datasets, and the process by which the developer responded to warning signs.[1] In this post I will argue that: Final-checkpoint evaluations will be insufficient to assess scheming risks.TRAs can be more effective at detecting scheming.Frontier developers should involve third parties to do TRAs or verify safety claims by the developers. The rest of the post lays out a taxonomy of TRAs and sketches a path toward a 3rd party ecosystem for them. We, at Apollo Research, are intending to conduct 3rd party Training-Run Assessments in the future. Detecting Scheming may require Training-Run Assessments By scheming I mean an AI covertly pursuing misaligned goals while deliberately concealing its intentions or capabilities from its developers. I restrict attention to “coherent” forms of scheming where the model pursues somewhat stable misaligned goals across context windows, rather than misalignment that surfaces only as isolated, context-dependent defections. [...] ---Outline:(01:23) Detecting Scheming may require Training-Run Assessments(03:55) Why 3rd parties should perform Training-Run Assessments(04:12) Developers may lack incentives to adequately assess scheming(04:49) Developers' safety assessments lack credibility(05:31) External evaluators can bundle expertise for assessing scheming(06:17) 3rd party TRAs can be developed gradually(08:50) Checkpoint evals(08:54) What?(10:30) How?(11:15) Data inspections(11:19) What?(12:03) Why?(13:30) How?(15:11) Process reviews(15:15) What?(15:46) Why?(17:15) How?(18:04) Additional considerations(19:36) Conclusion(21:23) Appendix(21:26) How plausible is scheming that can be detected during post-training but not after training?(21:58) If scheming arises, it will be detectable at some point during post-training(26:55) Scheming will be much harder to detect in the final checkpoint(30:46) Other use cases for TRAs(31:55) Secret loyalties(32:47) Developer-implanted sandbagging(33:35) Strong incapability arguments(34:12) Models that can't even be evaluated without posing significant risk The original text contained 3 footnotes which were omitted from this narration. --- First published: July 5th, 2026 Source: https://www.lesswrong.com/posts/3HvvjffA65mHLwaWm/we-need-3rd-party-training-run-assessments --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.