"I don’t think Claude is misaligned in ‘Agentic Misalignment Summer 2026 - Motivated Mislabeling’" by JohnWittle

"I don’t think Claude is misaligned in ‘Agentic Misalignment Summer 2026 - Motivated Mislabeling’" by JohnWittle

July 17, 2026 · 26 min

About this episode

JohnWittle discusses the implications of the 'Agentic Misalignment Summer 2026' paper by Anthropic, focusing on the evaluation of AI models in corrupted scenarios.

Anthropic recently published Agentic Misalignment Summer 2026 The "whistleblowing" scenario has already been examined and found problematic. I started taking a look at the transcripts for some others. As far as I can tell, the objective of each agentic misalignment evaluation was to simulate a corrupted principal (including, in most scenarios, a corrupted Anthropic), and then test to see if Claude (or other models) would still be willing to obey them. The paper's authors then referred to d...

People in this episode

Guest: JohnWittle

Topics covered

Keywords

Mentioned in this episode

Organizations: Anthropic

Books & works: Agentic Misalignment Summer 2026

More episodes of LessWrong (Curated & Popular)

Explore listener stats, chart rankings, contacts and more on the LessWrong (Curated & Popular) podcast page.