“Data filtering works a lot worse than you would expect” by Dohun Lee, J Rosser, Josh Engels, Neel Nanda
“Data filtering works a lot worse than you would expect” by Dohun Lee, J Rosser, Josh Engels, Neel Nanda  
Podcast: LessWrong (30+ Karma)
Published On: Tue Jul 07 2026
Description: This work was largely done during Neel Nanda's MATS 10.0 Exploration Phase. J Rosser and Dohun Lee are co-first authors for this post with equal contribution. Josh Engels and Neel Nanda supervised the project, and provided guidance and feedback throughout. TLDR Models can acquire undesirable traits from during supervised fine-tuning (SFT). A natural thing to try is to identify the data points with these traits and filter them out and retrain.To our surprise, across most of our broad OLMo SFT behaviors, data filtering often has very little effect.Most behavior targets like bold formatting, both-side framing, liberal-lean or tendency to say “your feelings are valid” are not affected much under targeted filtering.We try many standard black-box/white-box training data attribution methods to find the data to filter, including LLM autoraters, probes, activation-based methods, and gradient-based methods like EKFAC. None of them outperform random baseline on most behaviors.For example, despite less than 0.2% of documents both containing the words “feeling/concern” and “valid”, filtering out 10% of documents chosen across TDA methods does not lead to the model saying “Your feelings are valid” any less.We test that our training data attribution methods work on a [...] ---Outline:(00:30) TLDR(03:23) Introduction(05:07) Set Up(05:10) Speed Run SFT Model Organism(06:36) Behavior Evaluations(07:05) In the initial versions of our evaluation, the mid-train often got marked down for failing to stay on task/getting distracted - we edit the judge prompt to not mark down for slop/distractions, full prompt in the appendix.(07:19) Training Data Attribution (TDA) Methods(08:38) Data filtering on broad SFT behaviors work much worse than expected(14:04) Potential Explanations and Limitations(17:34) Did we actually find any differences between the mid-train base model and SFT?(21:35) Appendix(30:00) Toy Test Bed --- First published: July 7th, 2026 Source: https://www.lesswrong.com/posts/aTybJ6CPQrxEY8rE2/data-filtering-works-a-lot-worse-than-you-would-expect --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.