
The episode discusses various aspects of AI alignment and associated risks.
The Don’t Worry About the Vase Podcast is a listener-supported podcast. To receive new posts and support the cost of creation, consider becoming a free or paid subscriber. * 00:00 - Introduction * 01:15 - Table of Contents * 02:32 - Here We Go Again: Executive Summary * 03:58 - Introduction (1) * 04:03 - RSP Evaluations (2) * 05:11 - Move That Goalpost * 07:11 - The Failures Are News * 09:21 - Alignment Risk Slowly Rises * 10:52 - New Risk Pathways Just Dropped * 13:28 - Cyber (3) * 14:27 - Harmful Requests (4.1) * 16:46 - We Need To Talk (4.2 and 4.3) * 19:56 - Overcoming Bias (4.4) * 21:59 - Agentic Safety (5) * 24:38 - Prompt Injection (5.2) * 31:08 - Alignment (6) * 32:23 - Looking For Problems * 33:54 - Who Watches The Training (6.2.2) * 38:02 - Automated Behavioral Audit * 38:39 - The Model Is Smarter Than The Eval (6.2.3.2) * 40:48 - You Should See The Other Guy * 43:12 - UK AISI Testing (6.2.4) * 43:32 - In Vendbench (6.2.5) * 46:10 - Honesty (6.3.3 to 6.3.6) * 49:00 - Chain of Thought (CoT) Monitorability (6.5) * 51:46 - What’s In The Box? (6.6) * 54:01 - That’s All For Now…
Host: Zvi
Organizations: Don't Worry About the Vase Podcast
Explore listener stats, chart rankings, contacts and more on the Don't Worry About the Vase Podcast podcast page.