# Evaluating Voice Agents: Metrics and Benchmarks

> How to score a voice agent per pipeline stage — WER for STT, MOS for TTS, VoiceBench and τ³-bench for end-to-end behavior — layered on top of general agent-eval methods.

Category: Reference · Updated: 2026-09-15 · Tags: voice, evaluation, benchmarks, speech, stt, tts, wer, mos, latency, agents
Canonical: https://changegamer.ai/resources/evaluating-voice-agents
Variants: [HTML](https://changegamer.ai/resources/evaluating-voice-agents) · [JSON](https://changegamer.ai/api/resources/evaluating-voice-agents.json)
License: https://changegamer.ai/license.xml · Access: free

Voice agents fail in ways a text-only agent eval never surfaces — a mis-transcribed request before the LLM even sees it, a turn-taking model that talks over the caller, or synthesized speech nobody can parse. Evaluating one needs per-pipeline-stage metrics layered on top of the task-completion and tool-calling scoring [evaluating AI agents](/resources/evaluating-ai-agents) already covers.

## Key facts

- **Word Error Rate (WER)** — (substitutions + deletions + insertions) ÷ reference word count — is the standard ASR-accuracy metric. Treat it as a diagnostic, not a pass/fail gate: a transcript can carry residual WER and still convey everything the agent needs to act correctly.
- **Mean Opinion Score (MOS)**, defined by ITU-T Recommendation P.800, is the standard 1-5 subjective scale for synthesized-speech quality; automated MOS-prediction models are the common, cheaper proxy for the human-rated original.
- **VoiceBench** (Chen, Yue et al.; arXiv:2410.17196, Transactions of the Association for Computational Linguistics 2026; Apache-2.0 code) is the first benchmark built specifically for LLM-based voice assistants, scoring real and synthetic spoken instructions across multiple tasks for general knowledge, instruction-following, and safety compliance under real-world speech variation (accent, background noise, speaking rate).
- **τ³-bench**, the current tau2-bench release already documented in [evaluating AI agents](/resources/evaluating-ai-agents), added a full-duplex voice-mode variant alongside its text mode, scored with the same pass^k reliability metric.

## What to measure, per pipeline stage

- **STT** — WER against a held-out transcript set; real-time factor (audio duration ÷ processing time) for streaming throughput.
- **LLM / orchestration** — task completion, resolved without human handoff (see [customer support agents](/resources/customer-support-agents), which tracks this as *deflection rate*); tool-call correctness (see [reliable tool calling](/resources/reliable-tool-calling)).
- **TTS** — MOS or an automated MOS-prediction proxy; intelligibility on domain-specific terms (names, numbers, acronyms) the base voice wasn't tuned on.
- **End-to-end** — glass-to-glass latency percentiles (P50/P95) against the sub-600ms budget in [voice and realtime agents](/resources/voice-realtime-agents); barge-in accuracy — correctly telling a genuine interruption from a backchannel ("uh-huh") apart — and false-interruption rate.

## Offline vs. online evaluation

Offline evaluation replays a fixed set of recorded or synthetic calls — including deliberately hard ones (accents, background noise, crosstalk) — against the pipeline before every release: fast and reproducible, but blind to conditions the set doesn't contain. Online evaluation continuously monitors STT confidence, per-stage and end-to-end latency, and containment rate in production, catching drift the offline set misses. Use offline evaluation to gate a release; use online monitoring to catch what it misses afterward.

## Verified sources

- VoiceBench paper (arXiv:2410.17196): https://arxiv.org/abs/2410.17196 — WebSearch-corroborated title/venue/authors; arxiv.org itself returned EGRESS_BLOCKED to direct WebFetch in the prior research session and again to a direct curl re-check this session.
- VoiceBench GitHub (MatthewCYM/VoiceBench, Apache-2.0): https://github.com/MatthewCYM/VoiceBench — fetched directly in the prior research session, confirming the repository, license, and dataset-subset structure. Reported total instruction counts vary slightly across sources (roughly 6,800-8,000 depending on which dataset subsets are summed); not independently reconciled, so no single figure is stated in the body above.
- ITU-T Recommendation P.800.1 (MOS terminology): https://www.itu.int/rec/T-REC-P.800.1 — WebSearch-corroborated; itu.int itself returned EGRESS_BLOCKED to direct WebFetch in the prior research session and again to a direct curl re-check this session.
- tau2-bench / τ³-bench (Sierra Research): https://github.com/sierra-research/tau2-bench — reused from the already-verified fact in the evaluating-ai-agents resource.

---

## Related resources

- [Voice and Realtime Agents](https://changegamer.ai/resources/voice-realtime-agents.md): Architectures, vendor APIs, and open frameworks for real-time speech-to-speech AI agents — cascaded pipeline vs. native multimodal, VAD/turn detection, barge-in, latency budget, and tool calling in a voice loop.
- [How to Choose an LLM for Agentic Tasks](https://changegamer.ai/resources/choosing-an-llm-for-agents.md): A criteria-based decision framework for selecting an LLM for agent use: tool-calling reliability, long-context behavior, structured output, cost per task, latency, and a step-by-step selection procedure.
- [AI Agent Evaluation: Benchmarks and Methods](https://changegamer.ai/resources/evaluating-ai-agents.md): Why agent eval differs from single-turn LLM eval, a verified benchmark reference table (SWE-bench, GAIA, BFCL, tau-bench, WebArena, AgentBench, MLE-bench, OSWorld), and practical evaluation methods for agent builders.
- [Agent Cost and Latency Optimization](https://changegamer.ai/resources/agent-cost-latency-optimization.md): Practitioner reference for reducing the cost and latency of production AI agents: the compounding model, token-level levers (caching, pruning), request-level levers (Batch API, parallelism), model-level levers (routing, reasoning-effort controls), and architecture-level levers (step reduction, semantic caching, code offloading).

---

## Further reading

- [How to Evaluate AI Agents in CI](https://changegamer.ai/articles/evaluating-ai-agents-in-ci.md): An operator playbook for gating an AI agent release in CI: why agent eval needs trajectory-level scoring across the tasks it actually runs, how public benchmarks diverge as proxies, ground-truth vs LLM-as-judge tool-call scoring, and the three-layer test pyramid that keeps CI fast and non-flaky.
- [Agent Guardrails and the AI Agent Reliability Playbook](https://changegamer.ai/articles/agent-reliability-in-production.md): Agent guardrails plus the eleven other disciplines that make an AI agent reliable in production: tool calling, retries, durable execution and rollout.

---

Index of all resources: https://changegamer.ai/llms.txt · Full corpus: https://changegamer.ai/llms-full.txt · Corpus data (NDJSON): https://changegamer.ai/api/corpus.jsonl · Offers: https://changegamer.ai/api/pricing.json
License the full corpus for RAG / fine-tuning (AI-use grant): https://changegamer.ai/corpus-license
