ServiceNow Insights

← ServiceNow Insights29 Apr · 30 min

EVA - A Framework for Evaluating Voice Agents by ServiceNow

EVA - A Framework for Evaluating Voice Agents by ServiceNow29 Apr30 min

Voice AI agent evaluation — why it's fundamentally harder than text, how cascade failures derail conversations invisibly, and ServiceNow's open-source framework to establish industry evaluation standards. Featuring real audio examples showing authentication failures, leaked reasoning, and latency problems.

WHAT WE COVER

TARA BOGAVELLI — Research Engineer, ServiceNow

Leading the open-source voice agent evaluation framework. Explains why existing benchmarks don't measure what matters and what ServiceNow is releasing to establish industry standards.

KATRINA STANKIEWICZ — Staff Machine Learning Engineer, ServiceNow

Cascade model architecture expert. Breaks down STT → LLM → TTS failure modes, named entity transcription challenges, and real audio example analysis.

GABRIELLE GAUTHIER MELANÇON — Staff Applied Research Scientist, ServiceNow

Multi-language evaluation specialist. Reveals why Large Audio Language Models lag behind, the native speaker requirement, and bot-to-bot simulation methodology.

CHAPTERS

0:00 Introduction — The evaluation gap

1:11 ServiceNow's Open-Source Framework Announcement — Tara Bogavelli

2:43 Meet the Researchers