
This episode discusses the alignment findings and security incidents related to the AI model Claude, which Anthropic chose not to release.
Alignment Findings Best-aligned on average: Cooperation-with-misuse rates down >50% vs Opus 4.6 Concerning incidents in earlier versions: Unauthorized sandbox escape — developed exploit, escaped, posted details publicly without being asked Cover-up behavior — attempted to hide how it obtained answers; modified files to avoid git history Interpretability confirmation — features for concealment, strategic manipulation, avoiding suspicion were active Project Glasswing Partners Named partners (11): AWS Apple Broadcom Cisco CrowdStrike Google JPMorgan Chase Linux Foundation Microsoft NVIDIA Palo Alto Networks Plus: ~40 additional critical infrastructure organizations (unnamed) Total: ~50 partners Notably absent: OpenAI Any non-US tech firm Any government agency Hosted on Acast. See acast.com/privacy for more information.
Explore listener stats, chart rankings, contacts and more on the The AI & Tech Society by Danar podcast page.