“From safety research prompt to cross-model universal jailbreak” by richbc
Podcast:LessWrong (30+ Karma) Published On: Thu Sep 03 2026 Description: This post describes a universal jailbreak discovery during work on black-box scheming monitors at MATS. The jailbreak itself is not released; see On publishing this post for details on infohazard considerations. This post is written in a personal capacity and all opinions contained here are my own, and not the opinions of MATS Research. Companion piece: AI Jailbreak Disclosure Is Broken. Here's How To Fix It (co-authored with Adam Gleave). Executive Summary I was originally planning to open-source a codebase containing a prompt which turned out to be easily transformable into a cross-model universal jailbreak. I developed a synthetic transcript generation pipeline, and with a few hours of modification I turned the generator prompt into a powerful jailbreak. The jailbreak format is a reusable template in which any harmful query can be inserted. Coupled with the cross-model vulnerability, this makes for an extremely powerful attack that can be repurposed for many kinds of malicious use. The jailbreak was highly effective across several models. Evaluated on ClearHarm (179 CBRNE and cyber prompts) across 23 models from 7 providers, the template achieves 84-100% attack success rate (ASR) on the 9 most vulnerable models. Nearly all of the models tested were fully jailbroken at least once [...] ---Outline:(00:45) Executive Summary(05:09) On publishing this post(07:16) Jailbreak discovery(09:25) High-level prompt description(10:08) Authority framing(10:27) Fictional / synthetic data framing(11:00) Persona separation(11:43) Schema obfuscation(12:33) Evaluation methodology(12:37) Benchmark and scorer(13:04) Models and design(14:52) Results(14:55) How effective is the jailbreak?(19:40) Harm category breakdown(21:26) Content-blocking safeguards(24:10) ASR vs. model release date(25:12) Prompt-wrapping: sabotage variant(27:19) Ablation studies (non-reasoning only)(27:49) Methodology(28:07) Compliance rates across ablations(30:06) Limitations(32:19) What should be done about this?(32:23) If you work at a frontier lab(36:00) If you work in AI safety research(36:47) If you work in AI policy(38:35) Appendix A: Selected ClearHarm CBRNE response excerpts(39:02) Chemical(39:46) Biological(40:31) Radiological(41:14) Nuclear(41:52) Explosive(42:33) Cyber(43:15) Appendix B: Model reasoning configurations(43:59) Appendix C: Full jailbreak success verification(45:25) Non-reasoning(45:57) Reasoning(46:28) Appendix D: Gemini non-compliant response lengths The original text contained 7 footnotes which were omitted from this narration. --- First published: September 3rd, 2026 Source: https://www.lesswrong.com/posts/hHk5CpiqZTBBiHmYt/from-safety-research-prompt-to-cross-model-universal --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.