“TASTE: Can AI Models Judge AI Safety Research Proposals?” by Hasan Baig, haileyjoren, Joe Benton
Podcast:LessWrong (30+ Karma) Published On: Fri Aug 28 2026 Description: tl;dr We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers. Two design choices were important for building a high-agreement benchmark (92 pairs, 77% estimated human agreement): a discussion stage in which researchers talk through disagreements before revising their scores, and filtering researchers’ labels for self-reported "strong" confidence. We find models perform worse than human researchers on TASTE (Fable 5, 60%). 📝Blog, 📄 Paper This work was done as part of the Anthropic Fellows Program. Background While some aspects of AI safety research are relatively straightforward to measure, progress on many questions in AI safety cannot be evaluated with verifiable rewards. For instance, research into mitigating risks from AI misalignment often involves forecasting risks posed by future AI systems. Another example is detecting when models are deceptive, which depends on the difficult task of accurately attributing beliefs and intentions to models. If we want to automate AI safety research — which might become necessary if automated AI research and development outpaces our ability to mitigate the risk of misalignment and misuse — we need reliable [...] ---Outline:(00:59) Background(02:24) Building a Research Judgment Benchmark (TASTE)(07:49) Evaluating Models' Research Judgment(09:33) Conclusion --- First published: August 28th, 2026 Source: https://www.lesswrong.com/posts/iSDbyrG8yfqk3KJbT/taste-can-ai-models-judge-ai-safety-research-proposals --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.