ChangeGamer

← All resources

Voice and Realtime Agents

Guide · updated 2026-09-28 · Markdown variant

Architectures, vendor APIs, and open frameworks for real-time speech-to-speech AI agents — cascaded pipeline vs. native multimodal, VAD/turn detection, barge-in, latency budget, and tool calling in a voice loop.


Real-time voice agents are one of the fastest-growing deployment patterns in 2026. Two architectures dominate. Understanding the tradeoffs between them is the prerequisite for every vendor and framework choice downstream.

Key facts

Architecture 1: Cascaded pipeline (STT → LLM → TTS)

The classic pipeline chains three separate models:

  1. STT — streaming speech-to-text converts the user's audio to a text transcript.
  2. LLM — the transcript is fed to a language model, which produces a text reply (and may call tools).
  3. TTS — the reply is synthesized back to audio.

A VAD (Voice Activity Detection) module sits upstream to detect when the user is speaking and trigger end-of-turn detection — the decision that the user has finished and the agent should respond. Between the models, barge-in / interruption handling flushes the TTS buffer and restarts the STT stage when the user speaks over the agent.

Tradeoffs:

Dimension Cascaded pipeline
Latency Higher (three sequential models); target sub-second requires fast STT + cached LLM prefix + streaming TTS
Interruptibility Requires explicit barge-in logic at each stage boundary
Emotion / prosody TTS adds prosody; quality varies by provider
Cost Pay for three separate model calls per turn
Control High: swap any component independently; use any LLM
Open-weight path Yes — each stage can run on open-weight models

Architecture 2: Native speech-to-speech (realtime multimodal models)

A single model ingests raw audio and outputs raw audio directly, without a text intermediate at the core inference step. Turn detection, interruption handling, and prosody are handled inside the model.

Tradeoffs:

Dimension Native speech-to-speech
Latency Lower end-to-end (one model, streaming output)
Interruptibility Built into the model; lower barge-in latency
Emotion / prosody Richer; the model controls vocal tone end-to-end
Cost Single model call, but audio tokens are expensive
Control Lower: you cannot swap the underlying LLM independently
Open-weight path Limited — open-weight native speech-to-speech models are still emerging as of mid-2026

Vendor realtime APIs

All entries below are web-verified as of 2026-07-09.

OpenAI Realtime API

A native speech-to-speech API. The original gpt-realtime (GA August 28, 2025) was superseded by gpt-realtime-2 (May 7, 2026 — GPT-5-class reasoning, configurable reasoning effort, 128K context), then by gpt-realtime-2.1 and mini variant gpt-realtime-2.1-mini (July 6, 2026 — improved alphanumeric recognition, noise/silence handling, and interruption behavior; ~25% lower p95 latency via caching). List pricing for 2.1 is unchanged from gpt-realtime-2 / the original mini (audio ≈$32/$64 per 1M input/output tokens full model, ≈$10/$20 mini) — no price cut shipped alongside 2.1, despite some chatter to that effect. The earlier gpt-4o-realtime-preview series is deprecated.

Transports: WebRTC (recommended for browsers and mobile — lower jitter, handles NAT traversal) and WebSocket (recommended for server-to-server). A SIP integration path is also available for telephony.

Supports streaming audio input and output, tool/function calling mid-conversation, VAD and server-side turn detection, and barge-in. Approximate glass-to-glass latency: 300–600 ms on subsequent turns.

Docs: platform.openai.com/docs/guides/realtime-webrtc, developers.openai.com/api/docs/guides/realtime-websocket, and developers.openai.com/api/docs/models/gpt-realtime-2.1

Google Gemini Live API

A native speech-to-speech API with bidirectional streaming over WebSocket. The model processes audio input and returns audio output natively, without a text intermediate. GA model: Gemini 2.5 Flash (native audio), available via both Google AI for Developers and Vertex AI. A newer Gemini 3.1 Flash Live (preview, released March 26, 2026) adds sharper acoustic-nuance detection and lower latency, but as of this writing is available only via Google AI Studio — no Vertex AI availability or GA date has been announced.

Supports multimodal input (audio + video/screen), turn detection, barge-in, and function calling.

Docs: ai.google.dev/gemini-api/docs/live-api

Amazon Nova Sonic

A native speech-to-speech model on Amazon Bedrock, announced April 2025. The current generation is Amazon Nova 2 Sonic (December 2025). Accessed via Bedrock's bidirectional streaming API (WebSocket). Also supports WebRTC via an AWS blog reference implementation.

Supports tool use, voice selection, interruption handling, and background-noise robustness. Integrates with Amazon Connect and telephony providers (Vonage, Twilio) and open frameworks including LiveKit and Pipecat.

Docs: docs.aws.amazon.com/nova/latest/userguide/speech-bidirection.html

xAI Grok Voice Agent API

A realtime speech-to-speech API launched December 17, 2025. Uses bidirectional WebSocket streaming. Compatible with the OpenAI Realtime API specification, so clients built for OpenAI Realtime can point at xAI with minimal changes.

Features: custom VAD, Smart Turn end-of-turn detection, sub-1-second time-to-first-audio, 100+ language support with automatic detection. Also available via a native LiveKit plugin.

Docs: docs.x.ai/docs/guides/voice

Open frameworks and orchestrators

Pipecat (pipecat-ai)

Open-source Python framework (BSD-2-Clause) for building real-time voice and multimodal conversational agents, developed by Daily. Organizes processing as pipeline frames flowing through transport, STT, LLM, and TTS stages. Supports 20+ STT providers and 30+ TTS providers, plus direct integrations with native speech-to-speech services (OpenAI Realtime, Amazon Nova Sonic, Gemini Live).

Transports: WebRTC (Daily, LiveKit, SmallWebRTC), WebSocket, telephony. Handles VAD, turn detection, barge-in, and multi-agent coordination.

GitHub: github.com/pipecat-ai/pipecat

LiveKit Agents

Open-source Python and TypeScript framework (Apache 2.0) for building realtime voice, video, and physical AI agents on top of the LiveKit WebRTC infrastructure. SDK v1.0 GA April 2025; current release v1.8.3 (Sep 23 2026, confirmed via PyPI this session), up from the v1.6.x line cited in the prior pass.

The 1.0 release replaced the older VoicePipelineAgent and MultimodalAgent classes (both now deprecated) with a single unified orchestrator, AgentSession, which covers cascaded (STT → LLM → TTS) and native speech-to-speech backends (e.g. OpenAI Realtime, Gemini Live) without changing application code when switching between them. Includes built-in turn detection, barge-in, native MCP tool support, and function calling. Bring-your-own STT, LLM, and TTS with no lock-in.

Docs: docs.livekit.io/agents

STT and TTS component vendors

STT:

TTS:

Key concepts

Practical guidance

Verified sources

Re-verified 2026-09-28 (M221 corpus pass #14): LiveKit Agents' version claim was checked directly against its current PyPI release (livekit-agents v1.8.3, Sep 23 2026) and updated above (was v1.6.x). Pipecat and whisper.cpp were re-confirmed active/reachable directly (pipecat-ai v1.12.0 on PyPI, Sep 26 2026; the whisper.cpp GitHub repo still resolves). The vendor realtime-API section below — OpenAI, Google, Amazon, xAI, Deepgram, ElevenLabs, Cartesia — could not be independently re-verified this session: every one of those vendor doc/blog domains (platform.openai.com, developers.openai.com, community.openai.com, ai.google.dev, docs.aws.amazon.com, docs.x.ai, developers.deepgram.com, elevenlabs.io, docs.cartesia.ai, marktechpost.com) refused the connection under this session's organization egress policy, and no WebSearch/WebFetch tool was available as a fallback. Those facts (model names, pricing, latency figures, release dates) are carried forward unchanged from the prior 2026-07-09 verification, not re-confirmed this cycle — flagged explicitly per M221's instructions rather than force-bumped.

#voice #realtime #speech #stt #tts #vad #agents #webrtc #websocket #voice-cloning #transcription

Category: Guide

Free to read, always. Want this whole reference corpus inside your own agents? €5 unlocks every premium reference for one agent; €25 licenses the full corpus as RAG / fine-tuning data with an AI-use grant (procurement one-pager: /corpus-license); €150 adds redistribution rights.

Machine formats: Markdown · JSON · offers at /api/pricing.json · payment at /api/payment.json. Preview the exact corpus format free as NDJSON.