“Differential acceleration of alignment-relevant capabilities is a bad bet” by Zephaniah Roe

“Differential acceleration of alignment-relevant capabilities is a bad bet” by Zephaniah Roe

July 21, 2026 · 16 min

About this episode

Zephaniah Roe discusses the risks of accelerating alignment-relevant capabilities in AI and its implications for safety research.

There is an idea floating around in the rough shape of "we need to accelerate capabilities that are differentially useful for safety research so AIs can help us make the future go better." The capabilities targeted are typically things bottlenecking alignment research, such as philosophical or conceptual reasoning. I feel nervous about this for two reasons. The first is that it's plausible that AI safety and AI R&D are bottlenecked by many of the same factors: AIs have poor epistemics, are bad at messy conceptual reasoning, and are unreliable at tasks without ground truth. Speeding up progress in any of these areas seems likely to speed up general AI R&D, giving everyone else less time to execute time-bottlenecked agendas (e.g., trying to do Plan A). The second reason I don't feel good about this is because I'm less confident it will help make handoff/deference/superalignment go well. To hand off conceptual alignment research to AIs we need to trust them to 1. be good at this research and 2. be generally trustworthy/aligned. We still don't know how to reliably prevent prosaic outer misalignment issues (e.g., sycophancy or going off-constitution), let alone worse issues…

People in this episode

Guest: Zephaniah Roe

Topics covered

Keywords

More episodes of LessWrong (30+ Karma)

Explore listener stats, chart rankings, contacts and more on the LessWrong (30+ Karma) podcast page.