
This episode explores the engineering and design challenges of building production-grade voice AI systems at scale.
This episode is brought to you by the MLflow team. Check out more information at MLflow.org . What does it actually take to build voice AI at a billion-interaction scale? This episode features an ex-Amazon voice AI engineer who built customer support systems handling 2 billion+ interactions โ now working on next-gen voice agent platforms. Anurag digs deep into the real engineering tradeoffs, design patterns, and use cases that separate production-grade voice agents from demos. Voice Agent Use Cases // MLOps Podcast #372 with Anurag Beniwal, Member of the Technical Staff at ElevenLabs ๐๏ธ Topics covered: ๐น Cascaded vs. speech-to-speech โ Why cascaded systems still win in production, and how to make them feel natural without sacrificing control ๐น Latency masking โ Foreground/background model architecture and how to buy yourself time while deep retrieval runs ๐น Constellation of models โ Using Haiku for tool calling, fine-tuned smaller models for response generation, and why "one model for everything" breaks at scale ๐น Turn-taking & ASR challenges โ Why voice is harder than chat: accents, noise, silence detection, and domain-specific fine-tuning ๐น Level 1 vs Levelโฆ
Hyperbolic, MLflow
Explore listener stats, chart rankings, contacts and more on the MLOps.community podcast page.