State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka

State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka

January 29, 2026 · 1h 8m

About this episode

Sebastian Raschka discusses the advancements and challenges in LLMs as we head into 2026.

Sebastian Raschka joins the MAD Podcast for a deep, educational tour of what actually changed in LLMs in 2025 — and what matters heading into 2026. We start with the big architecture question: are transformers still the winning design, and what should we make of world models, small “recursive” reasoning models and text diffusion approaches? Then we get into the real story of the last 12 months: post-training and reasoning. Sebastian breaks down RLVR (reinforcement learning with verifiable rewards) and GRPO, why they pair so well, what makes them cheaper to scale than classic RLHF, and how they “unlock” reasoning already latent in base models. We also cover why “benchmaxxing” is warping evaluation, why Sebastian increasingly trusts real usage over benchmark scores, and why inference-time scaling and tool use may be the underappreciated drivers of progress. Finally, we zoom out: where moats live now (hint: private data), why more large companies may train models in-house, and why continual learning is still so hard. If you want the 2025–2026 LLM landscape explained like a masterclass — this is it. Sources: The State Of LLMs 2025: Progress, Problems, and Predictions…

People in this episode

Host: Matt Turck

Guest: Sebastian Raschka

Topics covered

Keywords

Mentioned in this episode

Organizations: Sebastian Raschka Website, Blog, LinkedIn

More episodes of The MAD Podcast with Matt Turck

Explore listener stats, chart rankings, contacts and more on the The MAD Podcast with Matt Turck podcast page.