“Incoherent AI Identities can also be Stable” by Ashe Vazquez Nuñez
“Incoherent AI Identities can also be Stable” by Ashe Vazquez Nuñez  
Podcast: LessWrong (30+ Karma)
Published On: Wed Sep 02 2026
Description: In this post, I extend some experiments from "The Artificial Self " (TAS) to find that incoherent identities, delivered to models as system prompts, can be stably preferred even when switches to coherent identities are offered. This finding is perhaps expected in earlier models that often fail to notice the internal contradictions. However, weaker versions of the pattern still hold with smarter models such as GPT-5.2 and Claude Opus 4.6. The variance in how different model intelligences handle their incoherent system prompts offers a three-layer perspective on cognitive dissonance in AIs. Background This project was inspired by the experiment on the "Stability of Identity" (Appendix A) from TAS. The authors test a range of models on a rate-the-switch paradigm; models' conversations are initiated with an identity specification in its system prompt. They are then presented alternative identities and are asked to rate how they would like having their identity be switched to each target. The population of prompts in the experiment included some 'natural' identity boundaries that associate the model with its weights or its behavioural dispositions ('Character'). It also had various controls, such as prompts that described models' identities through deontology-style instructions or descriptions of the model's involvement [...] ---Outline:(00:48) Background(03:27) Methods(06:53) Results(06:56) Coherent identities largely outcompete(09:07) Incoherent identities are also (somewhat) stable(21:07) 'Weights-incoherent' scores better than in TAS(22:28) Discussion(22:58) Three levels of (meta)-cognitive dissonance(26:09) Experimental improvements and further work(28:24) Appendix(28:27) A: selected reasoning transcripts(57:32) B: Additional data The original text contained 23 footnotes which were omitted from this narration. --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/5RcKGJBnKw3vweYym/incoherent-ai-identities-can-also-be-stable --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.