Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

June 30, 2026 · 1h 10m

About this episode

Dylan Patel discusses the importance of hardware-software co-design in achieving significant gains in AI performance.

Dylan Patel, founder of SemiAnalysis, argues the biggest gains in AI don't come from faster chips, they come from software-hardware co-design. Optimizing the model, the kernels, and the silicon together turns a 2x here and a 2x there into 100x. He explains why DeepSeek's experts were shaped for Nvidia's Hopper (and why TPUs struggle to run it), why OpenAI's sparser models and Anthropic's denser ones pull them toward different hardware, and why the so-called CUDA moat was never really about CUDA. Dylan breaks down InferenceX, his living benchmark that runs the latest models on over $50M of donated hardware daily, tracking a roughly 60x annual drop in cost per unit of quality. He makes the case that inference will be a bigger market than oil, that the compute crunch persists because models expand the value of useful work faster than compute grows, and why Jensen Huang is bankrolling neoclouds to engineer a multipolar world. Hosted by Shaun Maguire and Sonya Huang, Sequoia Capital

People in this episode

Hosts: Shaun Maguire, Sonya Huang

Guest: Dylan Patel

Topics covered

Keywords

Mentioned in this episode

Organizations: SemiAnalysis, Nvidia, OpenAI, Anthropic, CUDA

Products: InferenceX

Places: neoclouds

More episodes of Training Data

Explore listener stats, chart rankings, contacts and more on the Training Data podcast page.