
Hugo Bowne-Anderson interviews Sebastian Raschka about the future of LLM architecture and the importance of understanding AI models.
If you take a model release as an anchor point , let’s say Nemotron 3 or Qwen 3.5, you can go in both directions : You can either plug them into an agent and play around with that, or you can look, okay, what does the model look like under the hood ? What are the ingredients? What type of attention mechanism do they use? What are currently research techniques that could make that even better in the next generation of models? What can we swap out, basically? And I’m interested in both of these! Sebastian Raschka , Independent AI Researcher and author of Build a Large Language Model from Scratch , joins Hugo to talk about what’s changed in AI architecture, from post-training to hybrid models, and why understanding what’s under the hood matters more than ever for developers building in the agentic era . Sebastian’s upcoming book, Build a Reasoning Model from Scratch , currently available for pre-order on Amazon and in early access on Manning ! Vanishing Gradients is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber. We Discuss: * Ed Tech for Agents : should we design educational content specifically for agentic…
Host: Hugo Bowne-Anderson
Guest: Sebastian Raschka
Products: Build a Large Language Model from Scratch, Build a Reasoning Model from Scratch, Nemotron 3, Qwen 3.5
Explore listener stats, chart rankings, contacts and more on the Vanishing Gradients podcast page.