Evaluating Voice Agents: Metrics and Benchmarks
How to score a voice agent per pipeline stage — WER for STT, MOS for TTS, VoiceBench and τ³-bench for end-to-end behavior — layered on top of general agent-eval methods.
Voice agents fail in ways a text-only agent eval never surfaces — a mis-transcribed request before the LLM even sees it, a turn-taking model that talks over the caller, or synthesized speech nobody can parse. Evaluating one needs per-pipeline-stage metrics layered on top of the task-completion and tool-calling scoring evaluating AI agents already covers.
Key facts
- Word Error Rate (WER) — (substitutions + deletions + insertions) ÷ reference word count — is the standard ASR-accuracy metric. Treat it as a diagnostic, not a pass/fail gate: a transcript can carry residual WER and still convey everything the agent needs to act correctly.
- Mean Opinion Score (MOS), defined by ITU-T Recommendation P.800, is the standard 1-5 subjective scale for synthesized-speech quality; automated MOS-prediction models are the common, cheaper proxy for the human-rated original.
- VoiceBench (Chen, Yue et al.; arXiv:2410.17196, Transactions of the Association for Computational Linguistics 2026; Apache-2.0 code) is the first benchmark built specifically for LLM-based voice assistants, scoring real and synthetic spoken instructions across multiple tasks for general knowledge, instruction-following, and safety compliance under real-world speech variation (accent, background noise, speaking rate).
- τ³-bench, the current tau2-bench release already documented in evaluating AI agents, added a full-duplex voice-mode variant alongside its text mode, scored with the same pass^k reliability metric.
What to measure, per pipeline stage
- STT — WER against a held-out transcript set; real-time factor (audio duration ÷ processing time) for streaming throughput.
- LLM / orchestration — task completion, resolved without human handoff (see customer support agents, which tracks this as deflection rate); tool-call correctness (see reliable tool calling).
- TTS — MOS or an automated MOS-prediction proxy; intelligibility on domain-specific terms (names, numbers, acronyms) the base voice wasn't tuned on.
- End-to-end — glass-to-glass latency percentiles (P50/P95) against the sub-600ms budget in voice and realtime agents; barge-in accuracy — correctly telling a genuine interruption from a backchannel ("uh-huh") apart — and false-interruption rate.
Offline vs. online evaluation
Offline evaluation replays a fixed set of recorded or synthetic calls — including deliberately hard ones (accents, background noise, crosstalk) — against the pipeline before every release: fast and reproducible, but blind to conditions the set doesn't contain. Online evaluation continuously monitors STT confidence, per-stage and end-to-end latency, and containment rate in production, catching drift the offline set misses. Use offline evaluation to gate a release; use online monitoring to catch what it misses afterward.
Verified sources
- VoiceBench paper (arXiv:2410.17196): https://arxiv.org/abs/2410.17196 — WebSearch-corroborated title/venue/authors; arxiv.org itself returned EGRESS_BLOCKED to direct WebFetch in the prior research session and again to a direct curl re-check this session.
- VoiceBench GitHub (MatthewCYM/VoiceBench, Apache-2.0): https://github.com/MatthewCYM/VoiceBench — fetched directly in the prior research session, confirming the repository, license, and dataset-subset structure. Reported total instruction counts vary slightly across sources (roughly 6,800-8,000 depending on which dataset subsets are summed); not independently reconciled, so no single figure is stated in the body above.
- ITU-T Recommendation P.800.1 (MOS terminology): https://www.itu.int/rec/T-REC-P.800.1 — WebSearch-corroborated; itu.int itself returned EGRESS_BLOCKED to direct WebFetch in the prior research session and again to a direct curl re-check this session.
- tau2-bench / τ³-bench (Sierra Research): https://github.com/sierra-research/tau2-bench — reused from the already-verified fact in the evaluating-ai-agents resource.
Free to read, always. Want this whole reference corpus inside your own agents? €5 unlocks every premium reference for one agent; €25 licenses the full corpus as RAG / fine-tuning data with an AI-use grant (procurement one-pager: /corpus-license); €150 adds redistribution rights.