“Desiderata for functional welfare experiments on LLMs” by Rikhil Jhaveri, Jamie Johnson, David Africa
“Desiderata for functional welfare experiments on LLMs” by Rikhil Jhaveri, Jamie Johnson, David Africa  
Podcast: LessWrong (30+ Karma)
Published On: Mon Jul 06 2026
Description: TLDR LLMs appear to have functional welfare: coherent sets of behaviour that track how well things are going relative to their goals.  Improving model functional welfare matters for safety (low welfare may amplify misalignment) and for moral reasons (models may be or become moral patients). Naive interventions can fail in non-obvious ways. We argue any successful intervention must:  A) shift multiple welfare-constituting channels together and coherently. B) avoid corrupting the model's ability to register whether it is succeeding or failing. We survey various available interventions, and we find the two desiderata tend to trade off; we argue that synthetic document fine-tuning is the most promising. From this, we propose a concrete experiment: induce a low-welfare state (via the welfare vector derived in Han et al. 2026) and test whether SDF-instilled beliefs reverse pathological behaviours without harming goal monitoring.  S1) Introduction    As model behaviour becomes more complex, it has become increasingly useful to attribute functional mental states to models, such as beliefs, goals, and even emotion concepts. There is now also evidence that we can usefully attribute to them a notion of functional welfare: roughly, the set of dispositions and behaviours that express a model's representation of how well things are going for it relative to its goals. Gemma 3, for instance [...] ---Outline:(00:13) TLDR(01:23) S1) Introduction(04:40) S2) What would count as successfully changing a model's welfare?(05:01) 2.1) Measuring too narrowly(05:39) 2.2) Collateral damage to non-pathological components of welfare(08:41) 2.3) Summarising(09:27) S3) What interventions are available to us?(10:01) 3.1) System prompting(12:29) 3.2) Bias-augmented consistency training (BCT)(15:11) 3.3) Synthetic document fine-tuning (SDF)(16:49) 3.4) Activation-targeting intervention (steering, ACT)(18:52) 3.5) Other interventions considered(20:16) 3.6) Summarising(20:46) S4) SDF as a countermeasure to induced negative welfare(23:30) S5) Open questions The original text contained 8 footnotes which were omitted from this narration. --- First published: July 6th, 2026 Source: https://www.lesswrong.com/posts/yku6byxdeKREibe5L/desiderata-for-functional-welfare-experiments-on-llms --- Narrated by TYPE III AUDIO.