“Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits” by Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan_Panigrahi, srishti-git1110, Maxime Riché
“Don’t Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoors and Preserves Desired Traits” by Kajetan Dymkiewicz, Tim Farrelly, Adam Prada, Ishaan_Panigrahi, srishti-git1110, Maxime Riché  
Podcast: LessWrong (30+ Karma)
Published On: Sat Aug 08 2026
Description: Inoculation prompting (IP) aims to keep undesired traits in training data from becoming part of a model's default behaviour. IP applies the same inoculation prompt to all training examples and it leaves underspecified how the desired and undesired traits (DT and UT) should activate. Two failures follow: The model develops backdoors: conditional vulnerabilities through which prompts that do not directly request the UT can still elicit it. We call this UT leakage.The desired trait weakens under ordinary prompts. We introduce Stratified Inoculation Prompting (SIP), which uses diverse prompts on safe examples (DT-only). Compared to standard IP, SIP significantly reduces leakage and retains more of the desired trait. SIP requires little safe data: a 5% DT-only pool oversampled to 25% is enough for good performance. Figure 1. IP compared to SIP. In SIP, uncertain and contaminated examples keep the inoculation prompt as in standard IP, while confidently safe examples are oversampled and trained under prompts drawn from multiple non-eliciting control prompt categories. Since SIP relies on filtering examples, we test its robustness to classification errors and find a strong asymmetry: failing to inoculate examples containing the undesired trait reintroduces it, whereas unnecessarily inoculating safe examples is benign. This suggests [...] ---Outline:(06:16) Inoculation Prompting Underspecifies the Intended Conditionalisation(09:20) What successful conditionalisation looks like(10:14) Stratified Inoculation Prompting (SIP)(12:49) SIP reduces leakage while preserving the desired trait(14:14) Control prompts recover the desired trait, diverse prompts narrow the backdoor's activation boundary(16:42) Oversampling reduces the need for distinct safe data(18:54) SIP reduces Emergent Misalignment more than Uniform IP(20:13) Data filtering errors have an asymmetric impact(23:18) Limiting residual access to the undesired trait(23:50) Diluting the prompt-trait association(26:06) Password-locking the inoculation prompt(29:23) Limitations --- First published: August 7th, 2026 Source: https://www.lesswrong.com/posts/FS7GFsGsH7CSQLahy/don-t-inoculate-everything-stratified-inoculation-prompting --- Narrated by TYPE III AUDIO. ---Images from the article: u, examples from the same distinct subset are repeated. (A) Mean default UT expression under a neutral prompt. (B) Mean default DT expression. (C) Mean leakage across the six non-eliciting prompt families, excluding the inoculation prompt and explicitly eliciting requests. (D) Setting-specific leakage trajectories at u = 5%. Panels A–C average equally across all setups. The labelled references anchor the heatmap colour scales: SFT(DT+UT) and DT-only SFT in Panels A–B, and Uniform IP and DT-only SFT in Panel C. Increasing m while holding u at 1–5% generally reduces leakage, while producing smaller and less consistent changes in DT expression." style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.