“Differential acceleration of alignment-relevant capabilities is a bad bet” by Zephaniah Roe
Podcast:LessWrong (30+ Karma) Published On: Tue Jul 21 2026 Description: There is an idea floating around in the rough shape of "we need to accelerate capabilities that are differentially useful for safety research so AIs can help us make the future go better." The capabilities targeted are typically things bottlenecking alignment research, such as philosophical or conceptual reasoning. I feel nervous about this for two reasons. The first is that it's plausible that AI safety and AI R&D are bottlenecked by many of the same factors: AIs have poor epistemics, are bad at messy conceptual reasoning, and are unreliable at tasks without ground truth. Speeding up progress in any of these areas seems likely to speed up general AI R&D, giving everyone else less time to execute time-bottlenecked agendas (e.g., trying to do Plan A). The second reason I don't feel good about this is because I'm less confident it will help make handoff/deference/superalignment go well. To hand off conceptual alignment research to AIs we need to trust them to 1. be good at this research and 2. be generally trustworthy/aligned. We still don't know how to reliably prevent prosaic outer misalignment issues (e.g., sycophancy or going off-constitution), let alone worse issues that will make AIs less trustworthy in [...] ---Outline:(02:15) Examples of arguments for the acceleration of alignment-relevant capabilities(06:13) Why we should not do this kind of differential acceleration(06:18) Alignment bottlenecks are also capabilities bottlenecks(08:18) This could hurt time-bottlenecked strategies(09:48) I don't think that this would unblock superalignment/hand-off plans(11:53) But what about things like philosophy?(12:46) But isn't it pretty unlikely that you help labs make real capabilities progress?(14:03) My epistemic status The original text contained 4 footnotes which were omitted from this narration. --- First published: July 21st, 2026 Source: https://www.lesswrong.com/posts/FzGqnCkdKeTnZ9tjE/differential-acceleration-of-alignment-relevant-capabilities --- Narrated by TYPE III AUDIO.