{
  "slug": "evaluating-voice-agents",
  "title": "Evaluating Voice Agents: Metrics and Benchmarks",
  "description": "How to score a voice agent per pipeline stage — WER for STT, MOS for TTS, VoiceBench and τ³-bench for end-to-end behavior — layered on top of general agent-eval methods.",
  "category": "Reference",
  "tags": [
    "voice",
    "evaluation",
    "benchmarks",
    "speech",
    "stt",
    "tts",
    "wer",
    "mos",
    "latency",
    "agents"
  ],
  "updated": "2026-09-15",
  "premium": false,
  "rights": {
    "access": "free",
    "license": "https://changegamer.ai/license.xml",
    "pricing": "https://changegamer.ai/api/pricing.json",
    "payment": "https://changegamer.ai/api/payment.json"
  },
  "canonical": "https://changegamer.ai/resources/evaluating-voice-agents",
  "markdown": "https://changegamer.ai/resources/evaluating-voice-agents.md",
  "outline": [
    {
      "depth": 2,
      "text": "Key facts",
      "anchor": "key-facts"
    },
    {
      "depth": 2,
      "text": "What to measure, per pipeline stage",
      "anchor": "what-to-measure-per-pipeline-stage"
    },
    {
      "depth": 2,
      "text": "Offline vs. online evaluation",
      "anchor": "offline-vs-online-evaluation"
    },
    {
      "depth": 2,
      "text": "Verified sources",
      "anchor": "verified-sources"
    }
  ],
  "related": [
    {
      "slug": "voice-realtime-agents",
      "title": "Voice and Realtime Agents",
      "description": "Architectures, vendor APIs, and open frameworks for real-time speech-to-speech AI agents — cascaded pipeline vs. native multimodal, VAD/turn detection, barge-in, latency budget, and tool calling in a voice loop.",
      "url": "https://changegamer.ai/resources/voice-realtime-agents"
    },
    {
      "slug": "choosing-an-llm-for-agents",
      "title": "How to Choose an LLM for Agentic Tasks",
      "description": "A criteria-based decision framework for selecting an LLM for agent use: tool-calling reliability, long-context behavior, structured output, cost per task, latency, and a step-by-step selection procedure.",
      "url": "https://changegamer.ai/resources/choosing-an-llm-for-agents"
    },
    {
      "slug": "evaluating-ai-agents",
      "title": "AI Agent Evaluation: Benchmarks and Methods",
      "description": "Why agent eval differs from single-turn LLM eval, a verified benchmark reference table (SWE-bench, GAIA, BFCL, tau-bench, WebArena, AgentBench, MLE-bench, OSWorld), and practical evaluation methods for agent builders.",
      "url": "https://changegamer.ai/resources/evaluating-ai-agents"
    },
    {
      "slug": "agent-cost-latency-optimization",
      "title": "Agent Cost and Latency Optimization",
      "description": "Practitioner reference for reducing the cost and latency of production AI agents: the compounding model, token-level levers (caching, pruning), request-level levers (Batch API, parallelism), model-level levers (routing, reasoning-effort controls), and architecture-level levers (step reduction, semantic caching, code offloading).",
      "url": "https://changegamer.ai/resources/agent-cost-latency-optimization"
    }
  ],
  "furtherReading": [
    {
      "slug": "evaluating-ai-agents-in-ci",
      "title": "How to Evaluate AI Agents in CI",
      "description": "An operator playbook for gating an AI agent release in CI: why agent eval needs trajectory-level scoring across the tasks it actually runs, how public benchmarks diverge as proxies, ground-truth vs LLM-as-judge tool-call scoring, and the three-layer test pyramid that keeps CI fast and non-flaky.",
      "url": "https://changegamer.ai/articles/evaluating-ai-agents-in-ci"
    },
    {
      "slug": "agent-reliability-in-production",
      "title": "Agent Guardrails and the AI Agent Reliability Playbook",
      "description": "Agent guardrails plus the eleven other disciplines that make an AI agent reliable in production: tool calling, retries, durable execution and rollout.",
      "url": "https://changegamer.ai/articles/agent-reliability-in-production"
    }
  ],
  "body": "Voice agents fail in ways a text-only agent eval never surfaces — a mis-transcribed request before the LLM even sees it, a turn-taking model that talks over the caller, or synthesized speech nobody can parse. Evaluating one needs per-pipeline-stage metrics layered on top of the task-completion and tool-calling scoring [evaluating AI agents](/resources/evaluating-ai-agents) already covers.\n\n## Key facts\n\n- **Word Error Rate (WER)** — (substitutions + deletions + insertions) ÷ reference word count — is the standard ASR-accuracy metric. Treat it as a diagnostic, not a pass/fail gate: a transcript can carry residual WER and still convey everything the agent needs to act correctly.\n- **Mean Opinion Score (MOS)**, defined by ITU-T Recommendation P.800, is the standard 1-5 subjective scale for synthesized-speech quality; automated MOS-prediction models are the common, cheaper proxy for the human-rated original.\n- **VoiceBench** (Chen, Yue et al.; arXiv:2410.17196, Transactions of the Association for Computational Linguistics 2026; Apache-2.0 code) is the first benchmark built specifically for LLM-based voice assistants, scoring real and synthetic spoken instructions across multiple tasks for general knowledge, instruction-following, and safety compliance under real-world speech variation (accent, background noise, speaking rate).\n- **τ³-bench**, the current tau2-bench release already documented in [evaluating AI agents](/resources/evaluating-ai-agents), added a full-duplex voice-mode variant alongside its text mode, scored with the same pass^k reliability metric.\n\n## What to measure, per pipeline stage\n\n- **STT** — WER against a held-out transcript set; real-time factor (audio duration ÷ processing time) for streaming throughput.\n- **LLM / orchestration** — task completion, resolved without human handoff (see [customer support agents](/resources/customer-support-agents), which tracks this as *deflection rate*); tool-call correctness (see [reliable tool calling](/resources/reliable-tool-calling)).\n- **TTS** — MOS or an automated MOS-prediction proxy; intelligibility on domain-specific terms (names, numbers, acronyms) the base voice wasn't tuned on.\n- **End-to-end** — glass-to-glass latency percentiles (P50/P95) against the sub-600ms budget in [voice and realtime agents](/resources/voice-realtime-agents); barge-in accuracy — correctly telling a genuine interruption from a backchannel (\"uh-huh\") apart — and false-interruption rate.\n\n## Offline vs. online evaluation\n\nOffline evaluation replays a fixed set of recorded or synthetic calls — including deliberately hard ones (accents, background noise, crosstalk) — against the pipeline before every release: fast and reproducible, but blind to conditions the set doesn't contain. Online evaluation continuously monitors STT confidence, per-stage and end-to-end latency, and containment rate in production, catching drift the offline set misses. Use offline evaluation to gate a release; use online monitoring to catch what it misses afterward.\n\n## Verified sources\n\n- VoiceBench paper (arXiv:2410.17196): https://arxiv.org/abs/2410.17196 — WebSearch-corroborated title/venue/authors; arxiv.org itself returned EGRESS_BLOCKED to direct WebFetch in the prior research session and again to a direct curl re-check this session.\n- VoiceBench GitHub (MatthewCYM/VoiceBench, Apache-2.0): https://github.com/MatthewCYM/VoiceBench — fetched directly in the prior research session, confirming the repository, license, and dataset-subset structure. Reported total instruction counts vary slightly across sources (roughly 6,800-8,000 depending on which dataset subsets are summed); not independently reconciled, so no single figure is stated in the body above.\n- ITU-T Recommendation P.800.1 (MOS terminology): https://www.itu.int/rec/T-REC-P.800.1 — WebSearch-corroborated; itu.int itself returned EGRESS_BLOCKED to direct WebFetch in the prior research session and again to a direct curl re-check this session.\n- tau2-bench / τ³-bench (Sierra Research): https://github.com/sierra-research/tau2-bench — reused from the already-verified fact in the evaluating-ai-agents resource.",
  "sources": [
    "https://arxiv.org/abs/2410.17196",
    "https://github.com/MatthewCYM/VoiceBench",
    "https://www.itu.int/rec/T-REC-P.800.1",
    "https://github.com/sierra-research/tau2-bench"
  ]
}