He May Have Found AI's FEELINGS — Richard Ren, Center for AI Safety Researcher
Podcast:Doom Debates! Published On: Wed Sep 09 2026 Description: Richard Ren is a research engineer at the Center for AI Safety who graduated summa cum laude from the University of Pennsylvania. His new paper on AI wellbeing makes the bold claim that today’s AIs have measurable, human-like feelings, which the paper calls “functional wellbeing.” They act happy when they succeed and sad when they’re berated.We cover how you measure a language model’s happiness, the “AI drugs” his lab concocted, and why I count the findings as a Yudkowskian victory. Then we debate what it all means for AI consciousness, and whether AI is a moral patient. If we can’t rule out that AI models suffer, what do we owe them?Richard is careful never to claim the models are conscious, and he puts his P(Doom) at 50–65%, right alongside my 50%. The real disagreement is foxes vs. hedgehogs: he takes the data as it comes, while I say Yudkowsky’s theory called it twenty years ago. Enjoy the ride.Watch on YouTube: https://www.youtube.com/watch?v=1glFImnyp6oTimestamps00:00:00 — Cold Open00:01:12 — Introducing Richard Ren00:02:24 — What’s Your P(Doom)?™00:03:31 — From AI Skeptic to Safety Researcher00:08:16 — Why Care About AI Wellbeing?00:11:45 — AIs Have Coherent Utility Functions00:17:33 — Persona Selection00:19:56 — Foxes vs. Hedgehogs00:24:28 — How Coherent Are AI Preferences?00:27:42 — From Preferences to Wellbeing00:31:16 — Can You Trust an AI’s Self-Report?00:36:42 — What Makes AI Happy and Sad00:39:19 — The Liberal, College-Educated Persona Hypothesis00:44:29 — Debating Where AI Experiences Qualia00:52:06 — On Substrate Independence00:56:28 — Janus, Dysphorics, and Mind Crime00:58:58 — Which Images Make AIs Happiest?00:59:38 — Creating an AI Drug01:04:01 — AI Drugs Are Yudkowsky’s Paperclips01:05:47 — The Missing Mood01:06:50 — Does Richard Support Pausing AI?LinksRichard Ren on X (@notRichardRen) — https://x.com/notRichardRenCollision — Richard Ren's Substack — https://richardren.substack.com/"AI Wellbeing: Measuring and Improving the Functional Pleasure and Pain of AIs" — Richard Ren, Kunyang Li, Mantas Mazeika et al. (CAIS, 2026) — https://www.ai-wellbeing.org/"Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs" — Mantas Mazeika et al. (CAIS, 2025) — the coherent-preferences paper this work builds on — https://www.emergent-values.ai/"The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems" — Richard Ren et al. (2025) — the AI honesty benchmark — https://www.mask-benchmark.ai/"Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?" — Richard Ren et al. (NeurIPS 2024) — the meta-analysis of AI safety benchmarks — https://arxiv.org/abs/2407.21792"Representation Engineering: A Top-Down Approach to AI Transparency" — Andy Zou et al. (2023) — Richard's first CAIS collaboration — https://arxiv.org/abs/2310.01405Center for AI Safety — https://safe.ai/Center for AI Safety — careers — https://safe.ai/careersStatement on AI Risk (Center for AI Safety, May 2023) — organized by Dan Hendrycks — https://safe.ai/work/statement-on-ai-extinction-riskThe 2026 Singapore Consensus on Global AI Safety Research Priorities — https://aisafetypriorities.org/UK AI Security Institute (formerly the AI Safety Institute) — https://www.aisi.gov.uk/Sora — the OpenAI video model that blew up Richard's 30-to-50-year timeline three months after he wrote it down — https://en.wikipedia.org/wiki/Sora_(text-to-video_model)Google Gemini calls itself "a disgrace to my species" (Ars Technica, Aug 2025) — the self-deleting-AI anecdote — https://arstechnica.com/ai/2025/08/google-gemini-struggles-to-write-code-calls-itself-a-disgrace-to-my-species/Coherent decisions imply consistent utilities — Eliezer Yudkowsky (the coherence-theorems argument) — https://www.lesswrong.com/posts/RQpNHSiWaXTvDxt6R/coherent-decisions-imply-consistent-utilitiesSimulators — Janus's essay on LLMs as persona simulators — https://www.lesswrong.com/posts/vJFdjigzmcXMhNTsx/simulatorsJanus (@repligate) on X — https://x.com/repligateThe Hedgehog and the Fox — Isaiah Berlin's original essay — https://en.wikipedia.org/wiki/The_Hedgehog_and_the_FoxAI 2027 — https://ai-2027.com/Magnifica Humanitas — Pope Leo XIV's encyclical on AI (May 2026), the "AIs are not conscious" position — https://www.vatican.va/content/leo-xiv/en/encyclicals/documents/20260515-magnifica-humanitas.htmlImplicit Association Test (Harvard Project Implicit) — https://implicit.harvard.edu/implicit/PauseAI — https://pauseai.info/We Found AI's Preferences — Bombshell New Safety Research — I Explain It Better Than David Shapiro — https://www.youtube.com/watch?v=ml1JdiELQ30Gödel's Theorem Proves AI Lacks Consciousness?! Liron Reacts to Sir Roger Penrose — https://www.youtube.com/watch?v=xwvijjZxpwIHe Led The Famous 2023 Statement on AI Extinction Risk — Adam Khoja, Center for AI Safety Researcher — https://www.youtube.com/watch?v=QqESBXuo6EIDoom Debates’ Mission is to raise mainstream awareness of imminent extinction from AGI and build the social infrastructure for high-quality debate.Support the mission by subscribing to my Substack at DoomDebates.com and to youtube.com/@DoomDebates, or to really take things to the next level: Donate 🙏 Get full access to Doom Debates at lironshapira.substack.com/subscribe