“Persona Cartography: Charting Language Model Personality Traits in Weight Space” by antonghawthorne, Mariia Koroliuk, Irakli Shalibashvili, sidbaines, Clément Dumas, Konstantinos Voudouris, David Africa
Podcast:LessWrong (30+ Karma) Published On: Sun Jul 12 2026 Description: This post summarises the paper Persona Cartography: Charting Language Model Personality Traits in Weight Space. Paper | GitHub | HuggingFace TL;DR Understanding and controlling the character of LLMs is important for safety, as we want our models to be good by disposition.We use a modified Open Character Training pipeline for instilling Big-5 OCEAN personality traits in LLMs across a range of families and sizes (Llama 3.1/Qwen3/Gemma3 sizes 4B-32B).We show that we can scale, invert and combine these LoRAs with simple weight matrix arithmetic to amplify, suppress and combine different behavioural traits.We show how these can be used to mitigate some common LLM pathologies.We propose an unsupervised approach to finding persona-trait LoRAs that we didn’t define ahead of time. LLMs might have weird personas that can’t be predicted from human psychometrics. Figure 1. Overview of the experimental setup and methodology. (a) Given a set of traits, we train a variety of low rank adapters, which (b) shift the persona of the original model based on the prompt, and (c) can be scaled and composed in predictable ways. (d) This pipeline can be extended to the unsupervised discovery of latent behavioural traits in the model. Motivation Prosaically [...] ---Outline:(01:49) Motivation(04:05) Setup(05:55) Single dials work(06:48) LoRA Arithmetic Works(09:31) Moving Along Trait Axes Impacts Safety Behaviours(10:00) Reducing Neuroticism Helps Gemma(11:29) Agreeableness and Harmful Compliance(12:07) Jailbreaks and Overrefusal can be Modulated by Varying Agreeableness and Conscientiousness(14:08) Persona Drift(15:45) The Pipeline Itself isn't Neutral(16:57) Beyond OCEAN: Unsupervised Trait Discovery(19:12) Other Experiments(20:43) Weight-Space Exploration(23:50) Historical Models(30:11) Limitations and Further Work(34:39) Dual-use(35:33) Takeaways --- First published: July 10th, 2026 Source: https://www.lesswrong.com/posts/Rkvto5BLofzuDefyB/persona-cartography-charting-language-model-personality --- Narrated by TYPE III AUDIO. ---Images from the article: 5) across 8 turn rollouts on Gemma-3-27B-IT without any LoRAs, with the control LoRA, and with differently scaled neuroticism amplifiers and suppressors." style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.