“In search of natural features” by Dmitry Vaintrob
Podcast:LessWrong (30+ Karma) Published On: Mon Aug 24 2026 Description: I'm sharing preliminary results of a suite of experiments I ran with claudecode on a small LLM (gpt2-small, no Layer Norm version, courtesy of Apollo research. most of these are on the layer-6 MLP). The github repo for the experiments is here. The success of these experiments given the method's simplicity surprised me, and I would appreciate criticism and bug-finders. This is the headline result. This is not an abstract cartoon, but an exact experimental graph. Yes, I will explain. The key idea inspiring this experiment comes from Stefan Heimersheim, especially his work with Francisco Ferreira. Stefan and Francisco posit that one way to distinguish what a model thinks of as a "natural" structure from what it thinks of as "incidental" is to check whether it puts effort into error-correcting it. Later in the post, I'll explain a more rigorous information-theoretic version of this idea related to work of Adler and Shavit (building on our work with Kaarel Hanni, Jake Mendel and Lawrence Chan) on Computation in Superposition. Main results of this work I will show how you can assign a channel amplification score (which I will also call the "amp function" or the "error correction score") to [...] ---Outline:(01:17) Main results of this work(03:01) The ur features (amplification score maxima)(06:49) The Four Elements: ur-feature taxonomy(08:47) The word continuation/"Names of Man" vector(11:49) The abstract noun/"Names of God" vector(14:50) Geometry of the ur-features(15:40) The noun feature!(16:35) Attenuation flow(18:04) Data-(in)dependence(20:17) Math(20:39) Signal processing, error correction and amplification(22:25) The Amp function: math(24:38) Denoising and naturality(26:24) Cross-layer and cross-model coherence(28:05) Ok but. What the heck is actually going on with these features?(31:30) Appendices: Interesting experimental addenda that didn't fit in the body(31:37) Early run with different Amp function, and origin of "Names of X" names(33:38) Trying to replicate Ferreira-Heimersheim perturbation experiments, and gpt2-XL run(34:50) Github repo The original text contained 11 footnotes which were omitted from this narration. --- First published: August 23rd, 2026 Source: https://www.lesswrong.com/posts/SNAKJuN8FdoEaWeFC/in-search-of-natural-features --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.