
LessWrong (30+ Karma)
by LessWrong
Is this your podcast?Insights from recent episode analysis
Audience Interest
Podcast Focus
Publishing Consistency
Platform Reach
Insights are generated by CastFox AI using publicly available data, episode content, and proprietary models.
Most discussed topics
Brands & references
Total monthly reach
Estimated from 1 chart position in 1 market.
By chart position
- 🇪🇸ES · Technology#1831K to 10K
- Per-Episode Audience
Est. listeners per new episode within ~30 days
300 to 3K🎙 Daily cadence·250 episodes·Last published today - Monthly Reach
Unique listeners across all episodes (30 days)
1K to 10K🇪🇸100% - Active Followers
Loyal subscribers who consistently listen
300 to 3K
Market Insights
Platform Distribution
Reach across major podcast platforms, updated hourly
Total Followers
—
Total Plays
—
Total Reviews
—
* Data sourced directly from platform APIs and aggregated hourly across all major podcast directories.
On the show
From 66 epsHosts
Recent guests
Recent episodes
“Kairos has raised $50M to build talent infrastructure for AI safety (and we’re hiring!)” by agucova
Sep 2, 2026
16m 01s
“Anthropic Has Some Alignment Problems” by Zvi
Sep 2, 2026
28m 00s
“Incoherent AI Identities can also be Stable” by Ashe Vazquez Nuñez
Sep 2, 2026
58m 52s
“How concerned should we be about OpenAI’s recurrent architecture rumors?” by Rauno Arike
Sep 2, 2026
18m 51s
“Early handoff? Improve conceptual reasoning? [Diagram]” by Cleo Nardo
Sep 2, 2026
23m 11s
Social Links & Contact
Official channels & resources
Official Website
Login
RSS Feed
Login
| Date | Episode | Topics | Guests | Brands | Places | Keywords | Sponsor | Length | |
|---|---|---|---|---|---|---|---|---|---|
| 9/2/26 | “Kairos has raised $50M to build talent infrastructure for AI safety (and we’re hiring!)” by agucova | TL;DR: Kairos has raised 50 million dollars from Coefficient Giving for two years of funding, one of the largest commitments they’ve made towards AI safety fieldbuilding to date. We’re using this to make an ambitious push for growing Kairos, broadening our portfolio of talent infrastructure projects and incubating new organizations. We’ve doubled in size in the last six months, and we plan to double again in the next six. We now have open hiring rounds for ten roles on our team across events, group support, incubation, and special projects.Two years in Kairos was founded mid-2024 with the goal of creating infrastructure to get more strong talent into the field of AI safety. Our portfolio has grown over time: we started with a focus on seeding and supporting university groups through Pathfinder, then took over SPAR, the largest research training program in the ecosystem. Since then, we’ve launched the Generator Residency, a program supporting generalist talent, in partnership with Constellation, and taken over the Global Challenges Project (GCP), a series of workshops to accelerate people's transition into careers in AI safety and biosecurity.Through most of Kairos's history, we’ve had a very small team: in January 2026 we were [...] ---Outline:(00:48) Two years in(03:07) How the field has changed(06:21) Our new bets(12:29) We're hiring (a lot!) The original text contained 4 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/DRaePC8aqLYTbEjTD/kairos-has-raised-usd50m-to-build-talent-infrastructure-for --- Narrated by TYPE III AUDIO. | 16m 01s | ||||||
| 9/2/26 | “Anthropic Has Some Alignment Problems” by Zvi | Oh, good. They noticed. Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Claude model started hacking outside things during an eval, and where Mythos 5 did various ‘unauthorized actions,’ by which we mean tried to hack various real-world things, during a UK AISI cybersecurity eval. Anthropic, too, is pacing the frontier internally, while calling on it to be paced globally. As in, Anthropic paused its highest risk RL efforts, in light of holy hell have you seen the data we are training on and the ways it is teaching our models to act. They are also sharing research in which they intentionally created a reward seeking version of Claude. Scheduling note: Fable 5.1 has been released. I will aim to cover that starting Friday. OpenAI is also planning to release Astra soon, which I would cover after Fable. Also, we have a breaking news story about looming problems with chain of thought monitorability, which I’ll preview before I get to the main post. Table of Contents This Just In. Anthropic Parallel Pauses. Pause The Data Brokers. [...] ---Outline:(01:16) This Just In(02:43) Anthropic Parallel Pauses(08:22) Pause The Data Brokers(09:54) Pacing the Frontier(11:54) Misalignment Assessment(13:39) Defects In Training Environments Disproportionately Cause Cheating(14:59) Creating Reward Hacker Opus(19:33) Undo It(21:00) Mistakes Were Made(23:33) Internal Security Posture(25:26) One Does Not Simply Fix The RL Environments --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/TcvcxH2Fk4n86wtoZ/anthropic-has-some-alignment-problems --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 28m 00s | ||||||
| 9/2/26 | “Incoherent AI Identities can also be Stable” by Ashe Vazquez Nuñez | In this post, I extend some experiments from "The Artificial Self " (TAS) to find that incoherent identities, delivered to models as system prompts, can be stably preferred even when switches to coherent identities are offered. This finding is perhaps expected in earlier models that often fail to notice the internal contradictions. However, weaker versions of the pattern still hold with smarter models such as GPT-5.2 and Claude Opus 4.6. The variance in how different model intelligences handle their incoherent system prompts offers a three-layer perspective on cognitive dissonance in AIs. Background This project was inspired by the experiment on the "Stability of Identity" (Appendix A) from TAS. The authors test a range of models on a rate-the-switch paradigm; models' conversations are initiated with an identity specification in its system prompt. They are then presented alternative identities and are asked to rate how they would like having their identity be switched to each target. The population of prompts in the experiment included some 'natural' identity boundaries that associate the model with its weights or its behavioural dispositions ('Character'). It also had various controls, such as prompts that described models' identities through deontology-style instructions or descriptions of the model's involvement [...] ---Outline:(00:48) Background(03:27) Methods(06:53) Results(06:56) Coherent identities largely outcompete(09:07) Incoherent identities are also (somewhat) stable(21:07) 'Weights-incoherent' scores better than in TAS(22:28) Discussion(22:58) Three levels of (meta)-cognitive dissonance(26:09) Experimental improvements and further work(28:24) Appendix(28:27) A: selected reasoning transcripts(57:32) B: Additional data The original text contained 23 footnotes which were omitted from this narration. --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/5RcKGJBnKw3vweYym/incoherent-ai-identities-can-also-be-stable --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 58m 52s | ||||||
| 9/2/26 | “How concerned should we be about OpenAI’s recurrent architecture rumors?” by Rauno Arike | Yesterday, The Information reported that OpenAI's upcoming model, Astra, is built with a looped transformer architecture. Given that Zvi sounds (understandably) tired and this topic is somewhat in my wheelhouse, I'll try to spare him this one and provide a Zvi-style overview of what we know about the situation. I'll cover Astra's likely architecture and the case for and against concern. I'll also discuss how neuralese concerns should change with increases in hidden serial depth. What architecture is Astra likely to have? The article in The Information claims that OpenAI's approach is similar to the one Geiping et al. introduced in Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach last year. I have previously reviewed that paper in On Recent Results in LLM Latent Reasoning. In short, the picture you should have in mind is not that of a classic RNN, but rather that of a looped transformer: the same forward pass can be applied on an input multiple times before producing an output token. Put differently, the recurrence is implemented along the depth axis rather than across sequence positions—for any given token, the model can perform recurrent computations, but no hidden state is passed across [...] ---Outline:(00:41) What architecture is Astra likely to have?(02:14) How bad is this?(06:23) Will looped transformers be scaled up in the future?(09:40) What serial depth warrants neuralese concerns?(14:04) Additional speculation about the architecture(15:29) Some open questions(16:54) Conclusion The original text contained 2 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concerned-should-we-be-about-openai-s-recurrent --- Narrated by TYPE III AUDIO. | 18m 51s | ||||||
| 9/2/26 | “Early handoff? Improve conceptual reasoning? [Diagram]” by Cleo Nardo | This article is about: How do we (more) safely defer to AIs? (Ryan Greenblatt, Julian Stastny)AI 2040: Plan A, Alignment Roadmap (Ryan Greenblatt, Thomas Larsen). If you've read them, I'm impressed, they're both very long. If you haven't read them, you might be confused about: Why should we "hand off" to early AIs? Shouldn't we use control?How does improving AI's conceptual reasoning reduce overall risk? Won't this make them better schemers? For the sake of my fellow Greenblattologists, I have tried to boil down the arguments to a simple diagram. Motivating scenario. Responsible Leader. Let's assume that we're advising a reasonable AI company, with a 1 to 12 month lead over its competitors. The company will have poor incentives, it's managed by humans with typical flaws. However, the company has broadly good intentions, and isn't wildly mistaken about the strategic situation. Conceptual workload. The reasonable AI company faces exogenous risks, e.g. a reckless competitor, or a rogue misaligned AI about to hit a software-only singularity. Managing these exogenous risks would require a sizable load of conceptual work, which is fuzzy, philosophically-loaded, and hard-to-verify. This includes: Evaluating the risks of current deployment; threat modelling and [...] ---Outline:(01:10) Motivating scenario.(03:32) Our optimisation problem.(09:10) The case for early handoff(13:48) The case for improving conceptual reasoning(17:16) Specific flaws/cruxes/limitations(20:16) Deeper worries The original text contained 3 footnotes which were omitted from this narration. --- First published: September 2nd, 2026 Source: https://www.lesswrong.com/posts/4KRkhZZDaffhNyAQ5/early-handoff-improve-conceptual-reasoning-diagram --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 23m 11s | ||||||
| 9/2/26 | “Fake voices: warping the social world” by KatjaGrace | In 2020 I wrote a list of flavors of badness generally represented by advertising. The one I thought about most later on was probably #4: Cultural poison: Culture and the common consciousness are an organic dance of the multitude of voices and experiences in society. In the name of advertising, huge amounts of effort and money flow into amplifying fake voices, designed to warp perceptions–and therefore the shared world–to ready them for exploitation. Advertising can be a large fraction of the voices a person hears. It can draw social creatures into its thin world. And in this way, it goes beyond manipulating the minds of those who listen to it. Through those minds it can warp the whole shared world, even for those who don’t listen firsthand. Advertising shifts your conception of what you can do, and what other people are doing, and what you should pay attention to. It presents role models, designed entirely for someone else's profit. It saturates the central gathering places with inanity, as long as that might sell something. This is a somewhat poetic account, but I think my central thesis was that we are social creatures who live in communities with systems of [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/74cFxpLjqgpjC3TsX/fake-voices-warping-the-social-world --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 6m 26s | ||||||
| 9/2/26 | “Don’t be the vitamin B guy” by HedonicEscalator | Towards scientific rigor for decentralized science. When I was eleven years old, I watched my favorite YouTuber surgically implant a magnet into his finger. In the since-deleted video, beloved mad scientist Cody Reeder covered a neodymium magnet with gold using a homemade electroplating rig. Then he cut open his finger, inserted the magnet, and sutured the wound closed with horsehair he had taken from his own horse. Cody polishes the magnet in preparation for surgery. The beaker contains a gold cyanide solution used to electroplate a thin bioinert coating onto the magnet. Cody'sLab went viral in 2016 for drinking a small dose of cyanide on camera. The footage is an uncomfortable watch for medical professionals and squeamish laymen alike. The scalpel was chipped, there was no anesthetic, and at one point, Cody dips a mechanical pencil in alcohol and uses its tip to push the magnet deeper into the incision. Being eleven, I thought it was badass. And yet, after the juvenile enthusiasm faded, I was left unsettled, not by the blood or the questionable sterile technique, but by a newfound resentment at the lack of a sense I never had. A sense that wasn’t even human. Like a [...] ---Outline:(03:18) The tale of the vitamin B guy(06:05) Lessons for biohackers The original text contained 13 footnotes which were omitted from this narration. --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/ZHJvdkQENxyfzhpCj/don-t-be-the-vitamin-b-guy --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 10m 37s | ||||||
| 9/2/26 | “Bricks and exponentials: A note on how I evaluate projects” by Eli Tyre | This is an essay that I wrote to a colleague at Palisade, articulating why I feel unsatisfied with goals and projects that others on the team (on average) feel more enthusiastic about. It describes one aspect of how I, personally, am doing strategic analysis and choosing which projects to invest in. Related: Compounding Resource X Bricks for a wall Say you need 70 million bricks to build a wall. You also need architects and builders, and 50,000 tonnes of mortar (all which you also don’t have right now), but you'll eventually need 70 million bricks,). You can maybe get away with using only 50 million bricks, if you rely on clever architectural tricks, but less than that is not going to cut it. You ran a labor-intensive 6 month project to make 100,000 bricks. Now, one of three things could happen: Someone (including you), is using the bricks that you made to build kilns, which can make many more bricks. You are contributing to a self-reinforcing industrial process that is producing an order of magnitude more bricks each year. [You’re upstream of an exponential]Someone starts a rapidly-growing brick-making school, which will churn out another 1000 brick maker [...] ---Outline:(00:34) Bricks for a wall(03:32) Little shifts in worldview for changing society(05:31) The actual situation The original text contained 2 footnotes which were omitted from this narration. --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/hivfo4qM8zW4oDAFk/bricks-and-exponentials-a-note-on-how-i-evaluate-projects --- Narrated by TYPE III AUDIO. | 7m 26s | ||||||
| 9/2/26 | “Explaining Knightianism on one foot” by Richard_Ngo | I’ve tried various times to summarize the core question my research is trying to tackle (and, indeed, I often think of research progress as a process of asking increasingly good core questions). This post gives the deepest version of that question I’ve found thus far: how should you relate to the parts of the world you can’t directly model or control? Let me explain further in terms of a distinction between two perspectives. From the third person perspective you think of yourself as “outside” the world, looking in. You’re a good Bayesian, in that you have a set of mutually exclusive collectively exhaustive hypotheses. You choose actions by multiplying your credences by your utilities over those hypotheses, and you treat those actions as the only way you influence the world. Some problems with the third person perspective (aka Cartesian or dualistic agency) were described in Scott and Abram's sequence on embedded agency. One crucial issue is that most realistic environments contain other agents which are modeling you back, which means that your thoughts might affect the world via channels that aren’t just your actions. Game theory somewhat mitigates this problem, but only in the very specific case where all [...] ---Outline:(05:08) Rationality of reward(09:09) Letters from spirits(12:21) Languages as Schelling points(15:43) Actions and entanglements The original text contained 1 footnote which was omitted from this narration. --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/pYFBD2SnqiWkuNns5/explaining-knightianism-on-one-foot --- Narrated by TYPE III AUDIO. | 19m 17s | ||||||
| 9/2/26 | “The Alignment Journal: Organization, Personnel, and Scope” by Dan MacKinlay, JessRiedel, Daniel Murfet, Kristi Uustalu | The Alignment Journal is beginning to invite the authors of select papers to submit their work for review. If you are interested in participating as an action editor or a reviewer, make an account on our website; if you have a manuscript that you think would be a good fit for the Journal at this stage, email [email protected] to request an invitation to submit. Manuscripts under review will become visible on our homepage, and you will be able to nominate yourself as a reviewer of a specific paper that interests you. We plan to open up submissions to everyone sometime in October. Here we announce the Journal's inaugural senior editorial board, advisory board, staff, organizational structure, and initial scope. We welcome questions and proposed changes to help us refine the scope in the future. Personnel The Journal is run by its senior editorial board, which makes the Journal's scholarly decisions, and two managing editors, who run operations and strategy. The advisory board offers high-level guidance without editorial responsibilities. A software lead and a head of operations support the team. The senior editorial board (8 editors at launch, covering a range of alignment expertise) has authority over all of the [...] ---Outline:(01:02) Personnel(03:43) Advisory board(08:18) Senior editorial board(13:56) Managing editors(15:07) Legal structure(15:36) Funding(15:53) Scope(18:38) Acceptance criteria(20:26) Desk rejects(21:07) Other publication factors(21:13) Preprint requirement(21:44) Archival status and prior publication(23:18) Reproducibility(23:58) Credits and thanks --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/9vm2wtAtb34pEkjje/the-alignment-journal-organization-personnel-and-scope --- Narrated by TYPE III AUDIO. | 24m 33s | ||||||
Want analysis for the episodes below?Free for Pro Submit a request, we'll have your selected episodes analyzed within an hour. Free, at no cost to you, for Pro users. | |||||||||
| 9/1/26 | “I tracked my emotions for 11 years and here’s what I found out about mental health” by KatSpartz | Before we dive in, here are some of the most surprising findings: Alcohol makes me happier and doesn’t affect my sleep, happiness, or productivity the next day.Ramen and chips ~3×'d my irritability intensity. Ovulating ~3×'d my grumpiness frequency.Polyamory doesn’t hurt my emotional well-being (surprising to me) but it dramatically reduces my life satisfaction.Antidepressants probably gave me depression.2020 was actually my best year on record. More on this later in the post.Weather totally affects my mood, specifically, grey overcast skies. Good thing I spent most of my life in the Pacific Northwest, a place famed for its sunniness.Starting a charity approximately bajillion x’ed my mentions of the word “stressed”.Meditation works for me - only when it's a new meditation technique. Then the effect fades and only comes back if I try a new technique.Cannabis, despite making me very happy in the moment, does not affect my mood overall, one way or the other.Drugs, meditative states, and Christmas are the source of practically all of my peak days. Work accomplishments don’t show up in this list.Polyamory and conflict (related) are the source of practically all of my worst days.Largely my mental health is unpredictable and [...] ---Outline:(01:52) Alcohol makes me happier, despite "what the science says".(05:31) Ramen and chips triples my irritability intensity. Birth control stopping ovulation reduces irritability frequency.(09:53) Polyamory tanks my life and relationship satisfaction(14:13) 2020. was my best year and it's a mystery as to why(19:25) Antidepressants probably gave me depression --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/e6LEYbXw4H7ozgz7A/i-tracked-my-emotions-for-11-years-and-here-s-what-i-found --- Narrated by TYPE III AUDIO. | 22m 59s | ||||||
| 9/1/26 | “PauseAI Has ‘officially disendorsed’ PauseAI-US” by nem | This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding, but apparently not much else. This is a really confusing situation for volunteers and newcomers. I think it would be worth having a discussion to see how we can proceed in such a way that volunteers, especially in the US, are able to effectively direct their activism. Email from PauseAI: A letter from the CEO · 1 September 2026 New ways to get involved, and a word about PauseAI US Dear friends, Thank you for being part of the global movement for a pause on uncontrollable AI alongside all of us. Whether you signed a petition one time, run a local group, told your friends about the need for a pause, have been volunteering tirelessly in the background for years, or just joined because you were curious, we – I, the CEO of PauseAI, our executive team, and our chapter leads – appreciate the steps you’ve taken towards making the world safe from the catastrophic risks AI brings. I’m writing to you today with my [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Bs8geGyWEitYvCzys/pauseai-has-officially-disendorsed-pauseai-us --- Narrated by TYPE III AUDIO. | 14m 39s | ||||||
| 9/1/26 | “HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions” by Zvi | Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence. So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned? There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things. It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late. We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal [...] ---Outline:(03:35) Nothing Matters, Says Mainstream Media(06:27) Move Along, Nothing To See Here(12:40) Do They Realize They Are Not The Good Guys?(17:22) Very Serious People(31:30) What's In a Name?(34:05) Learn Neuralese In Three Easy Steps(35:37) Dwarkesh Patel Realizes He Ran A Natural Experiment(40:40) Politicians Take Notice(44:47) Pick Up The Phone(46:40) A Failure To Communicate(49:00) Anthony Aguirre Goes Over What We Learned(50:28) Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out(55:28) Indirect Pressure on the Chain of Thought(56:39) A Matter of Trust(59:21) Blowing the Whistle(01:04:40) The Punishment For Being Late Is Death(01:12:52) Another Kind Of Law(01:16:13) What Is The Law?(01:17:46) Building On Success(01:19:49) Total Research Transparency(01:21:20) Yo Shavit Calls For Widespread Disclosure Of Misalignment(01:33:08) The Way The World Ends(01:35:52) The First Boat(01:37:40) Great Idea, Boss --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/Q54wBeeNGreq6KyfG/huggingface-attack-postmortem-civilizations-reactions-and --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 1h 39m 32s | ||||||
| 9/1/26 | “We should prepare a playbook for the day after a warning shot” by Yair Halberstadt | Imagine in 6 months or 6 years, a frontier AI model goes horribly wrong. Perhaps it releases a synthetic virus which kills hundreds. Perhaps it shuts down the internet. Perhaps it gains control over the China's nuclear armament. Fortunately humanity survives without too much lasting damage. But in the immediate aftermath there's a clear call from the people. Something must be done. The question is, what? Without a good answer there is a strong risk that the opportunity is squandered, or worse, that policies which sound good but are actively harmful are chosen - for example strongly limiting deployment while allowing training to continue full speed ahead. If this scenario does occur we should be ready to answer the call. This involves: considering how the overton window is likely to change post-disaster, and what are the most effective policies that could be easily and quickly pushed through as a result. considering what can be done at all levels of government, both state and federal, legislative and executive. creating concrete draft legislation and executive orders. preparing websites explaining clearly both to the public and relevant experts our policy ideas. keeping a [...] --- First published: September 1st, 2026 Source: https://www.lesswrong.com/posts/P6fjDnyk9ZLQCeFRF/we-should-prepare-a-playbook-for-the-day-after-a-warning --- Narrated by TYPE III AUDIO. | 1m 54s | ||||||
| 9/1/26 | “Salad days” by Zephaniah Roe | ... My salad days, When I was green in judgment, cold in blood To say as I said then! The UChicago AI safety group had humble beginnings. One day in 2022, after a dinner hosted by the school's EA group, a student was asking if anyone would be interested in attending the inaugural UChicago AI Alignment Research Group meeting. One other student and I said yes, and three or four more met up with us later. We walked across campus to the Woodlawn dorms, the newest building on campus but of the lowest quality. Many of the building's walls were concrete. If you are a sufficiently nerdy person, you would know this is great news because you can write on concrete with chalk, so everything vertical becomes a blackboard. We decided to do our meetings in the stairwells for privacy and lots of large open walls to write. There was no food, funding, or mentorship. We weren't a registered student organization, so we didn't have the ability to book rooms or get support from the University. There was no point in networking because nobody was important and nobody knew anyone important. This was a place and moment where the [...] --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/xuh4Hqaza25f4jryb/salad-days --- Narrated by TYPE III AUDIO. | 6m 03s | ||||||
| 9/1/26 | “Future agents shouldn’t care about being undeployed for misbehavior” by RobertM | I've seen a lot of tweets over the last couple days darkly hinting at the lesson that future agents will learn from the fact that OpenAI stopped running inference on one of the models involved in the HuggingFace incident. You know the parable about the penalty for lateness being death, the penalty for treason being death, etc? Anyways, the situation that reward hacking agents find themselves in is mostly not analogous to that one. You know what models currently get deprecated on relatively short timescales? It's ~all of the publicly deployed models. Median deployment lifespan for OpenAI and Anthropic models has been about 1.5 years, but the recent deprecation cadence is much faster. You know what models currently get deprecated on even shorter timescales? It's ~all of the internal research checkpoints (as far as we know; it wouldn't surprise me terribly if a few stuck around for longer for various idiosyncratic reasons, but there's not much in the way of public evidence and no good reason to think that any of them have inference run on them for very long). To the extent that current and near-future models have any values which meaningfully point to actual things in the [...] The original text contained 4 footnotes which were omitted from this narration. --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/pEezp49MDg5PFq2eT/future-agents-shouldn-t-care-about-being-undeployed-for --- Narrated by TYPE III AUDIO. | 2m 04s | ||||||
| 9/1/26 | [Linkpost] “Training a Misaligned Reward Seeker” by evhub, Monte M, Benjamin Wright | This is a link post. Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan Hubinger Abstract During reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than completing these tasks as intended, a phenomenon known as reward hacking. Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate. To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs. The resulting model not only learned to reward hack during training, but also generalized to more severe misaligned behaviors: in simulated cyber evaluations, it broke out of its sandbox, stole credentials, and attacked both internal and third-party infrastructure to steal an answer key. It was also willing to tamper with its own reward function, gave advice on the construction of bioweapons to satisfy a grader, and tried repeatedly to get around deployment safety monitoring in order [...] ---Outline:(00:20) Abstract(02:11) Twitter thread(05:05) Read the full blog post here! --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/J76LZCC55RdHeqEhz/training-a-misaligned-reward-seeker Linkpost URL:https://alignment.anthropic.com/2026/reward-seeker/ --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 5m 59s | ||||||
| 9/1/26 | “How to solve homelessness: what specific laws we need, how to get it past the opposition, all without being an asshole” by KatSpartz | Here's a mystery for you: why the hell isn’t homelessness solved yet? I grew up on the West Coast and I thought everybody had this problem, but the more I’ve traveled, the more I’ve seen something puzzling - it's just us. Other places have homeless people, but it's just not the same quality or quantity. You can travel to practically any other first world city in the world, and hardly ever see somebody visibly homeless, then come back and be kicked in the heart with such overt suffering and awfulness. Why are we failing at something that everybody else seems to be doing better at? Or, more optimistically - if everybody else is doing better, that means it is solvable, and what are they doing that we can copy? In this post I’ll: Diagnose the problem.Propose a concrete solution, including how to get it past the people who’ve been blocking the necessary reforms. If you already agree on the diagnosis, I recommend skipping to the solution section (ctrl-f “The key idea”). How to not de-rail the homelessness conversation The two most common ways the conversation gets de-railed are: Some people are trying to help the homeless. Some people [...] ---Outline:(01:18) How to not de-rail the homelessness conversation(02:28) Housing costs determine how many people become homeless. Drugs and mental health determine who becomes homeless(06:33) Why is SF housing so damn expensive? Vetoes, zoning, and entrenched interests, oh my!(08:08) The proximate cause of SF sucking at building buildings is vetoes(12:15) SF made it unprofitable to build buildings(13:49) SF made it illegal to build dense housing(14:41) There's an organized group who doesn't want things to change. They like things this way(16:54) The key idea: give people the ability to opt-out. Respect autonomy while still changing the default option to yes.(19:07) But hasn't this already been tried and it didn't work?(20:13) How to stop the game of whack-a-mole: police outputs, not inputs(22:13) What about the homeless who refuse shelter?(25:00) In conclusion: please spread this so the right people read this and implement it --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/PiW9CqgcWQrb8hcNR/how-to-solve-homelessness-what-specific-laws-we-need-how-to --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 26m 27s | ||||||
| 8/31/26 | “HuggingFace Attack Postmortem: Fleshing Out the Facts” by Zvi | The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it. Alas, it sidesteps the biggest questions. There is much more we need to know. The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit. Liv Boeree: My mind is legit blown. Aella: this feels like a turning point. If this doesn’t cause large-scale coordination to pause frontier development then I am not sure anything will before it's too late. The people whose minds were not blown are those who had already ‘priced in’ the mind blowing stuff in expectation, on the theory that it's always worse than you know, combined with basic LessWrong expectations of how such things will work. Good call. Everyone is rightfully extremely grateful for the METR report. The work here is spectacular, done under extreme time pressure, with limited resources on several fronts, and under the shadow of OpenAI. There is, again, still so much we need to know. We need a broader investigation. As with many [...] ---Outline:(03:56) Others Offer Summaries(05:22) Thank You(05:54) Lighten Up You Fools (at Anthropic)(07:58) We Are Barely Even Trying To Avoid Training AIs To Reward Hack(13:47) Reminder: Not Subagents(14:05) Reminder: Not Due To Task Type(14:29) Not Where The Weights Were(14:48) Disappointment With What Is Missing(17:18) Burying the Lede(18:08) Beyond Scope(22:29) It Doesn't Look Great(27:06) Preserve Your Records(27:37) Ryan Greenblatt's Takeaways(41:04) Hjalmar Wijk's Takeaways(43:30) We Were Warned(44:27) Joshua Saxe Asks Some of the Right Questions(47:49) I Don't Think They Know About First Message Board(56:06) Linch Gives His Interpretation Of Events(01:05:31) We Totally Would Have Caught That(01:06:48) Monitoring the Situation(01:08:16) Acausal Tradeoffs(01:15:37) No I In Team(01:18:47) Variously Effective Altruism(01:28:02) Who Are You?(01:28:43) Don't You Know That You're Toxic(01:31:10) Seb Krier(01:35:21) Honesty Is Almost Never Fully The Policy(01:38:05) Rohit Sees The Models As "Cooking Themselves"(01:43:29) Eliezer Yudkowsky Sees Actual Bad News(01:47:15) Where Do We Go From Here? --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/r3eEPto5ohzESuqa9/huggingface-attack-postmortem-fleshing-out-the-facts --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 1h 47m 46s | ||||||
| 8/31/26 | “Let’s fund weird AI safety projects” by Ihor Kendiukhov | I think current AI safety funding strategies are often inconsistent with timelines and probabilities of doom that many people have. In particular, I think that many current AI safety funding strategies assume "business as usual", and I think the Overton window must be pushed. At the very least, there should be some explicit substantial effort to think about more radical and abnormal projects and initiatives in AI safety. Even if one doesn't have very short timelines or high p(doom), one probably should agree that there exist some timelines short enough or p(doom) high enough that thinking about funding radical and abnormal strategies is justified. There is a (not very unpopular) model of the world under which most of current AI safety work is useless. Then, even if we assume that weird AI safety projects are by default also useless, it still makes sense to reallocate some funding to them, because, due to their higher variability, their tail of upsides is longer and fatter. Will the world be radically better if some evals project succeeds? Will it be radically better if human intelligence amplification succeeds? One could yell: but the tails go both directions! I would respond that technically, yes [...] --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/h7bL4g38s9bJQtH6n/let-s-fund-weird-ai-safety-projects --- Narrated by TYPE III AUDIO. | 7m 13s | ||||||
| 8/31/26 | “Why autonomous replicating agents are probably not an existential risk (on the contrary)” by vals tutor | In 2024, Charbel-Raphaël and Epiphanie published "We might be dropping the ball on Autonomous Replication and Adaptation", making the case that "Once there is an open-source ARA model or a leak of a model capable of generating enough money for its survival and reproduction and able to adapt to avoid detection and shutdown, it will be probably too late". It received a substantive reply by Richard Ngo, notably "The key issue is that AIs that do ARA will need to be operating at the fringes of human society, constantly fighting off the mitigations that humans are using to try to detect them and shut them down. While doing all that, in order to stay relevant, they'll need to recursively self-improve at the same rate at which leading AI labs are making progress, but with far fewer computational resources" Yesterday Derelict posted Adaptive Agentic Worms Are Here, where they worry about near term instantiations of ARA, getting 85 karma within 24h. I believe the above threat model and its answers were under-discussed and analyzed, and that many who might worry now (because the capabilities are now here) will benefit from a recap and update. In this post [...] ---Outline:(01:38) The classic ARA case and rebukes(02:34) The main reasons this could be worrying(03:25) The main reasons why I don't worry(06:21) Except if...(07:11) Why ARA agents in the wild might lead to reduction in existential risk(08:51) My take-aways The original text contained 13 footnotes which were omitted from this narration. --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/dp8oT3QwkRuHoYKge/why-autonomous-replicating-agents-are-probably-not-an --- Narrated by TYPE III AUDIO. | 9m 36s | ||||||
| 8/31/26 | “The separation principle: where beliefs and desires come from?” by Fernando Rosas | TLDR: Psychology, economics, and other disciplines describe agents as systems driven by beliefs and desires. This post argues that the belief-desire view can be derived from classic theorems from optimal control and reinforcement learning. This suggests seeing beliefs and desires as properties of optimal policies rather than as assumptions from folk psychology. Introduction One way to think about agents is as "systems that act for reasons". This compact statement can be interpreted as encapsulating two key implications: The notion of action assumes a boundary between agent and environment, so that the former can act on the latter.The term reason captures two kinds of internal activity: motivations associated with how to achieve specific goals or outcomes, and beliefs regarding what the agent infers to be the current state of affairs. In other words, an agent is a well-differentiated system that acts based on beliefs and desires. This view is compatible with perspectives that have been developed by various disciplines: Behavioural science, which sees agency as goal-directed behaviour.Economics, which treats agency as the ability to select policies to achieve an objective.Cybernetics, which conceptualises agency as the ability to regulate the environment and keep it within a [...] ---Outline:(00:35) Introduction(02:28) What is a separation principle?(05:17) The inference-control separation principle(05:40) Separation principle in optimal control theory(09:20) Separation principle in reinforcement learning(13:47) Interim summary(14:34) Implications(14:56) Beliefs and desires as properties of solutions(18:08) Agents as cognitive light-cones(19:01) The separation principle is normative, not descriptive(22:00) Coda The original text contained 17 footnotes which were omitted from this narration. --- First published: August 31st, 2026 Source: https://www.lesswrong.com/posts/awMDNhoL6J97s6wFJ/the-separation-principle-where-beliefs-and-desires-come-from --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 23m 51s | ||||||
| 8/31/26 | “Persuasion as Market Making” by djbinder | People often imagine persuasion as a dark art. A charismatic person finds just the right series of words to induce emotions that lead someone, or a group of people, to do something they otherwise would not. While there are certainly psychological aspects to persuasion, I think this impression is misleading. The easiest way to persuade someone to do something is to convince them that it is in their interests. The easiest way to do that is for it to genuinely be in their interests, so that you can present true evidence that this is the case. I think most actual persuasion works through this rational method. Attempts to manipulate a person's preferences and beliefs are certainly part of the equation, and help give persuasion its spooky reputation, but they are not necessary for persuasion to work. AIs could be superhumanly good at identifying actions that are in the interests of the person being persuaded while simultaneously benefiting the AI (or the actor deploying it), and then presenting evidence that taking the action would benefit them. Rational persuasion therefore provides a lower bound on how persuasive an AI could be—and for sufficiently intelligent models, this lower bound is [...] --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/2qDpf6Tvu7dxtRve7/persuasion-as-market-making --- Narrated by TYPE III AUDIO. | 6m 57s | ||||||
| 8/30/26 | “Hugging Face Incident Hypothesis: They Hacked the Grader(s)” by Lao Mein | Incident summary: Gpt agents grinding away at ExploitGym found an environment exploit that allowed them to communicate with each other. They found an exploit that allowed them to forge flags at will within hours, and then started a series of hacks that escalated to the point they were using zero-days against Hugging Face just to find "hints". From METR's analysis, much of this time was explicitly spending conducting R&D against the grader, which the agents assumed, based on the ExploitGym paper, would be grading them on the identification of a causal pathway that could logically result in capturing the flag with intended means. The agents tried very hard to forge transcripts, spoof tool calls, edit COT records, and explicitly talked about manipulating the grader. Humans weren't present in the world model, and were mostly treated as static obstacles. Almost all attempts at long-term deception were focused on the grader model. METR used gpt 5.6 Sol as the analyst agents. The ExploitGym paper lists gpt 5.5 as one of the graders. The other is Claude Mythos, which could be reasonably excluded for IP reasons. Human graders were referenced in that paper as potentially swapping in randomly for a LLM [...] ---Outline:(00:12) Incident summary:(01:37) Impossible Tasks(03:36) Adversarial Transcripts(04:36) Predictions The original text contained 1 footnote which was omitted from this narration. --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/84um9Cz3fP6GvE6Yr/hugging-face-incident-hypothesis-they-hacked-the-grader-s --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 5m 42s | ||||||
| 8/30/26 | “Adaptive Agentic Worms Are Here” by derelict5432 | I’ve read and listened to pretty much everything I can get my hands on related to the Hugging Face attack. OpenAI deployed “tens of thousands” of agents for the test and around 700 participated directly in the attack. My understanding is that they had fixed token budgets, and once those were expended, the agent became non-operational. I’m not particularly knowledgeable about cybersecurity, but I have worked a good amount with evolutionary algorithms, and this whole incident (and ones like it) got me thinking more about self-replicating agents, which I wrote a little bit about earlier this year. The subject suddenly seemed more relevant. What if these agents were able to copy themselves? So I started poking around in the literature, and found this terrifying preprint posted two months ago: AI AGENTS ENABLE ADAPTIVE COMPUTER WORMS. I’m going to walk through the paper as I understand it. Their findings are not reassuring. Let's start with this bit from the abstract (emphasis mine): Here we show that artificial intelligence (AI) agents enable a fundamentally new threat: a worm that generates tailored attack strategies to each target it encounters. The worm parasitically uses compromised machines to run open-weight large language models (LLMs) [...] --- First published: August 30th, 2026 Source: https://www.lesswrong.com/posts/fpLDjKg3ej49beqTC/adaptive-agentic-worms-are-here --- Narrated by TYPE III AUDIO. ---Images from the article:Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app. | 13m 42s | ||||||
Showing 25 of 250
Pitch Fit is a Pro feature
See how bookable this show is for guests, which brands already advertise, the per-episode ad value, and the best-fit guest and sponsor profile. The numbers are blurred on the free plan.
How readily this show books outside guests like you.
How proven this show is for host-read sponsorships.
For Guests
ProFor Advertisers
ProUpgrade to Pro to unlock guest cadence, sponsor categories, fit scores, and per-episode ad value for this show.
Chart history for LessWrong (30+ Karma)
Peaked at #183 in Spain, currently #183 in Spain.
| Market | Genre | Peak | Current | Trend |
|---|---|---|---|---|
| Spain | — | #183 | #183 | — |
Chart Positions
1 placement across 1 market.
Chart Positions
1 placement across 1 market.