
Stephen Balaban discusses the misconceptions around GPU compute and the emerging neoclouds in AI infrastructure.
Many people said GPU compute would become a commodity. The opposite happened — and a new category of "neoclouds" is now racing to build the physical backbone of the AI boom. Stephen Balaban, co-founder and CTO of Lambda, explains why the conventional wisdom was exactly wrong, why we're still massively underbuilding compute, and what it actually takes to stand up a gigawatt-scale AI factory: land, power, cooling, networking, and a financing stack most people have never heard of. We go deep on the physics of how energy becomes tokens, NVIDIA's real moat, why a 2023 GPU can lease for more today than the day it shipped, and Stephen's provocative vision of "neural software." Plus the wild Lambda origin story — from a facial recognition startup to a camera in a baseball cap to a near-billion-dollar cloud business. This is the state of AI compute in 2026, from inside one of the companies building it. (00:00) — Cold open (01:21) — Why GPU compute was never a commodity (02:45) — The H100 price index and what it gets wrong (04:02) — The real moat: technology or financing? (05:57) — Winner-take-all, or room for many neoclouds? (06:48) — Are we overbuilding or underbuilding AI compute…
Host: Matt Turck
Guest: Stephen Balaban
Organizations: Lambda
Products: H100
Explore listener stats, chart rankings, contacts and more on the The MAD Podcast with Matt Turck podcast page.