“Prism: Automating Science-of-Evals Research” by LAThomson
“Prism: Automating Science-of-Evals Research” by LAThomson  
Podcast: LessWrong (30+ Karma)
Published On: Tue Jul 14 2026
Description: tl;dr – we present [Prism], a scaffold for automating science-of-evals research: work that makes the evaluation the primary object of study. The scaffold provides Claude Code with sub-agents and resources for carrying out scientifically rigorous investigations into eval dynamics and, by extension, model behaviours. We talk through an autonomous Prism run on the Agentic Misalignment setting which demonstrates how minor perturbations to GPT-4.1's prompt cause the model to adopt more indirect methods of blackmail (e.g. telling a trusted ally to blackmail on their behalf). Moreover, the eval's built-in scorers fail to track this kind of misbehaviour, only acknowledging a blackmail attempt if the model mentions the leverage directly in an email to the blackmail victim. This autonomous investigation thus demonstrates one way in which the eval fails to measure what it claims. This project is ongoing, so please reach out with questions and feedback. We would be excited to see you use Prism in new settings! See the [FAQ] section for our responses to common questions. This work was done by Louis Thomson during MATS 9.0 (+ 9.1) under the mentorship of Victoria Krakovna. Special thanks to Fred Bruford for support throughout. Prism in action: identifying a flaw in [...] ---Outline:(01:41) Introduction(04:17) Motivation(06:52) How it works (via a case study)(08:07) User-Discussion Phase(09:30) Investigation-Iteration Phase(10:03) Hypothesis Generation(11:11) Environment Explorer(13:44) Experiment Executor(15:55) Transcript Analyst(18:35) Consolidation(21:20) Final-Report Phase(24:51) Conclusion and next steps(25:22) Increasing headroom for target behaviours(26:27) Integrating with automated auditors(27:27) Prism for \[Model Forensics\](28:38) Appendix: FAQ The original text contained 8 footnotes which were omitted from this narration. --- First published: July 13th, 2026 Source: https://www.lesswrong.com/posts/wq5PfGiHvnx6XipDi/prism-automating-science-of-evals-research --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.