“Modular Pretraining Enables Access Control” by E.Roland, cloud
Podcast:LessWrong (30+ Karma) Published On: Thu Jul 09 2026 Description: Full author list: Ethan Roland*, Murat Cubuktepe*, Erick Martinez*, Stijn Servaes, Keenan Pepper, Mike Vaiana, Diogo Schwerz de Lucena, Judd Rosenblatt, Addie Foote, Cem Anil, Alex Cloud; *Equal contribution tldr: Frontier AI models have knowledge that could be misused for nefarious purposes. To address this risk, we introduce Gradient Routed Auxiliary Modules (GRAM), a method for isolating dangerous knowledge to specific modules within a language model. These modules can be switched on or off to control what the model knows, making it possible to restrict or extend access to the most sensitive model capabilities based on user need and trust. In our experiments, we find evidence that a single model trained in this way can approximate multiple models, each trained with a different category of dangerous data filtered out, and this ability holds for models ranging from 50M to 5 billion parameters. This research is preliminary and has not been applied to production models at Anthropic. 📄 Paper, 💻 Code, 🖥️ Site This work was done at AE Studio, in collaboration with Anthropic. This is a cross-post from the Anthropic Alignment science blog. Introduction One of the major threats from frontier AI models is the misuse of legitimately helpful [...] ---Outline:(01:30) Introduction(05:30) The Method(07:48) Results(07:51) GRAM Matches Data Filtering(09:54) Access Control on Real Dual Use Data(12:23) Modularization Works Across Scales(14:05) Advantages of GRAM(14:15) Composability(15:30) Isolation Under Partial Labeling(17:18) Discussion(19:44) Acknowledgements --- First published: July 8th, 2026 Source: https://www.lesswrong.com/posts/43vKjWuH4goLwrFHA/modular-pretraining-enables-access-control --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.