“A proposal for a highly effective AI safety org” by ceselder
Podcast:LessWrong (30+ Karma) Published On: Thu Sep 03 2026 Description: TLDR: an org that pays people to "just read the fucking transcripts"; a large amount of people reading anonymized claude code/RL/eval transcripts flagged by a very high recall low precision monitor could catch warning shots, reward hacking and general weird stuff without needing to absorb any good people and this doesn't seem to exist. 5 million dollars per month could pay 1000 people to process literally all tokens in a frontier RL run and would catch ~15 serious incidents per month in bearish estimates. "Can't the models do it": I expect humans to remain necessary for the slice of "monitors don't catch it" and "is obvious egregious misalignment". (I'm thinking of things similar in nature to the message board stuff). This slice is obviously very important.In the limit, there's also scary inner alignment/scheming stuff that makes me want to have humans on this. Napkin math More concretely: We imagine ~100 people that are paid to use something like docent to read traces that are flagged as sus by a very low FNR monitor (or at the beginning, literally monte carlo sampling). Low hanging fruit is cyber/bio/behavioural evals/RL would be first but it could expand to more suspected benign [...] --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/FbKbfXjHvmZd8wtrX/a-proposal-for-a-highly-effective-ai-safety-org --- Narrated by TYPE III AUDIO.