ChangeGamer

← All resources

Evaluating Voice Agents: Metrics and Benchmarks

Reference · updated 2026-09-15 · Markdown variant

How to score a voice agent per pipeline stage — WER for STT, MOS for TTS, VoiceBench and τ³-bench for end-to-end behavior — layered on top of general agent-eval methods.


Voice agents fail in ways a text-only agent eval never surfaces — a mis-transcribed request before the LLM even sees it, a turn-taking model that talks over the caller, or synthesized speech nobody can parse. Evaluating one needs per-pipeline-stage metrics layered on top of the task-completion and tool-calling scoring evaluating AI agents already covers.

Key facts

What to measure, per pipeline stage

Offline vs. online evaluation

Offline evaluation replays a fixed set of recorded or synthetic calls — including deliberately hard ones (accents, background noise, crosstalk) — against the pipeline before every release: fast and reproducible, but blind to conditions the set doesn't contain. Online evaluation continuously monitors STT confidence, per-stage and end-to-end latency, and containment rate in production, catching drift the offline set misses. Use offline evaluation to gate a release; use online monitoring to catch what it misses afterward.

Verified sources

#voice #evaluation #benchmarks #speech #stt #tts #wer #mos #latency #agents

Category: Reference

Free to read, always. Want this whole reference corpus inside your own agents? €5 unlocks every premium reference for one agent; €25 licenses the full corpus as RAG / fine-tuning data with an AI-use grant (procurement one-pager: /corpus-license); €150 adds redistribution rights.

Machine formats: Markdown · JSON · offers at /api/pricing.json · payment at /api/payment.json. Preview the exact corpus format free as NDJSON.