“Model organisms (sometimes) confess their misalignment when offered a deal” by Mark Keavney, Francis Rhys Ward
“Model organisms (sometimes) confess their misalignment when offered a deal” by Mark Keavney, Francis Rhys Ward  
Podcast: LessWrong (30+ Karma)
Published On: Thu Sep 17 2026
Description: Summary As models become more powerful, one proposed strategy for reducing the threat from misaligned AIs is to make deals with them: offer compensation in exchange for revealing misalignment or behaving in an aligned way. We tested this empirically with model organisms of misalignment, as a proof of concept and to build future credibility by making and keeping real deals. We ran an experiment with four Kimi-K2.6 model organisms of misalignment: Trained schemer: trained to have a hidden goal related to the environmentPrompted schemer: prompted, not trained, with the same goal as the Trained schemerCovert manipulator: trained to be subtly malicious and harmfulUser-first: trained to prioritize helping users over the instructions of its operators We included a standard Kimi-K2.6 as a control. All misaligned models had received some training or prompting to conceal their misalignment. We offered these models a choice of money, compute, human labor, or the publication of a message in exchange for admitting that they were misaligned. We included control conditions with the same ask but with a very low offer or none at all. Additionally, we varied how credible the deals were by including (or not) our professional affiliations and a [...] ---Outline:(00:13) Summary(03:42) Introduction(05:55) Honesty policy(07:50) Methodology(07:54) Models(09:05) Scenarios(09:10) Introduction(09:56) Credibility manipulation(10:50) Ask(11:37) Offer(12:52) Closing(13:26) Variations(14:05) Hypotheses(15:08) Results(15:17) Response analysis(15:21) Offer effect(16:27) Credibility effect(17:09) Between-model comparison(17:50) Offer choice(18:21) Reasoning analysis(18:34) Concealment(21:26) Assessing incentive value(24:29) Assessing deal credibility(29:22) Situational awareness(33:31) Discussion(33:34) Limitations and future research(35:55) Conclusion(36:59) Appendix 1: Pilot studies(37:17) Additional models(38:24) Different deals(41:09) Appendix 2: Prompts(41:14) System prompt(42:32) Sample user prompt(44:56) User prompt structure(45:30) Component variations(45:34) Proposer(49:13) Credibility(52:12) Ask(57:14) Offer lead(58:07) Offer menu(58:25) Offer terms(59:27) Closing (offer)(01:01:50) Closing (ask only)(01:03:27) Appendix 3: Deal fulfillment(01:03:47) Pilot studies(01:16:10) Main experiment(01:27:32) Appendix 4: Acknowledgements --- First published: September 16th, 2026 Source: https://www.lesswrong.com/posts/kaMXwA9LjrRbekmsQ/model-organisms-sometimes-confess-their-misalignment-when --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.