
The episode discusses the challenges of evaluating voice AI agents and introduces ServiceNow's open-source framework for establishing industry standards.
Voice AI agent evaluation — why it's fundamentally harder than text, how cascade failures derail conversations invisibly, and ServiceNow's open-source framework to establish industry evaluation standards. Featuring real audio examples showing authentication failures, leaked reasoning, and latency problems. WHAT WE COVER TARA BOGAVELLI — Research Engineer, ServiceNow Leading the open-source voice agent evaluation framework. Explains why existing benchmarks don't measure what matters and what ServiceNow is releasing to establish industry standards. KATRINA STANKIEWICZ — Staff Machine Learning Engineer, ServiceNow Cascade model architecture expert. Breaks down STT → LLM → TTS failure modes, named entity transcription challenges, and real audio example analysis. GABRIELLE GAUTHIER MELANÇON — Staff Applied Research Scientist, ServiceNow Multi-language evaluation specialist. Reveals why Large Audio Language Models lag behind, the native speaker requirement, and bot-to-bot simulation methodology. CHAPTERS 0:00 Introduction — The evaluation gap 1:11 ServiceNow's Open-Source Framework Announcement — Tara Bogavelli 2:43 Meet…
Explore listener stats, chart rankings, contacts and more on the ServiceNow Insights podcast page.