ChangeGamer

← All guides · Agent observability and evaluation

AI Agent Observability and the Production Evaluation Playbook

Pillar guide · 4,236 words · ~19 min read · published 2026-09-19 · updated 2026-10-03 · Markdown variant

AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.

In short

  • AI agent observability and AI agent evaluation are two distinct disciplines built on the same trace data: observability captures what an agent actually did on a given run, and evaluation judges whether what it did was any good.
  • A retrieval-attribution log records which specific documents or context chunks a retrieval step pulled into an agent's working context, letting an operator tell a retrieval failure apart from a generation failure instead of guessing between the two.
  • A live judge model screening production AI agent output carries the same position, verbosity, and self-preference biases documented for offline evaluation, which is why a production system typically routes its low-confidence or high-stakes flags to a human queue rather than trusting every verdict unreviewed.
  • Cost telemetry for a production AI agent means capturing what each span and trace actually spent in tokens and dollars, a measurement discipline distinct from the separate question of how to reduce that spend.
  • An incident postmortem for an AI agent becomes durable only once the specific input or trajectory that caused the failure is added as a new fixture to the same evaluation suite that gates future releases, turning one incident into a permanent check.
  • No resource in this site's corpus defines a standard "escalation rate" metric for AI agents as of September 2026, so a team tracking one should set its own baseline rather than borrow a number from elsewhere.

An agent that fails loudly on day one is a cheaper problem than one that quietly does the wrong thing for three days before anyone notices — observability is what lets an operator tell the difference, and evaluation is what tells them whether the difference matters. This guide surveys both disciplines as they operate on an AI agent that is already live: tracing spans for tool calls and retrieval steps, screening live output with a judge model versus routing it to a human, the online metrics worth watching once real traffic replaces a benchmark set, what a running agent actually costs, and how an incident turns into a fixture the next release has to pass.

Two things this guide deliberately does not do. It does not re-derive pre-ship, CI-gating evaluation methodology — the public benchmark landscape, pass@k versus pass^k, the three-tier CI test pyramid — which how to evaluate AI agents in CI already covers in full; this guide treats evaluation only at the depth needed to place it next to observability, and points there for anything that gates a release before it ships. It also does not re-derive trace and span field mechanics — the specific gen_ai.* OpenTelemetry attribute names, or the five-field-per-span checklist — which what should an AI agent's observability system capture already covers in full; this guide treats tracing only at the depth needed to make the topic legible, and points there for the wire-level detail.

The rest of this guide moves through six disciplines in roughly the order a live agent actually produces the data they need: tracing, retrieval attribution, judge-versus-human screening, online metrics, cost telemetry, and postmortem-to-fixture feedback.

What is AI agent observability, and how does it differ from evaluation?

AI agent observability is the practice of capturing what an agent actually did on a given run — its trace of model calls, tool calls, and retrievals — while AI agent evaluation is the separate practice of judging whether what it did was correct, safe, or good enough. Observability answers "what happened": a trace ties one full run under a single ID, and each span inside it nests a model call, tool call, retrieval step, or sub-agent hop, as agent observability and tracing documents in full. Evaluation answers "was it any good": scoring a run, or a batch of runs, against a rubric, a benchmark, or a reference answer, as evaluating AI agents documents in full.

The table below is a map of the six disciplines this guide covers and how they relate, not a substitute for reading each section — it exists so a reader can place a specific problem (a cost spike, a bad live answer, a repeat incident) against the discipline that actually owns it before diving into the detail:

Discipline Question it answers What feeds it
Tracing What did the agent actually do on this one run? Nested spans under a shared trace ID
Retrieval attribution What context produced this specific output? Retrieval-step spans, kept distinct from other tool spans
Judge/human screening Was this specific live output any good? A sampled slice of live traffic plus a rubric
Online metrics Is behavior drifting across many runs, not just one? Feedback, behavioral signals, escalation share
Cost telemetry What did this run or trace actually spend? Per-span token counts converted to cost
Incident postmortem What should never happen the same way twice? A frozen trace from a run that already failed

Every row after the first two depends on the trace data the first two produce — which is why tracing and retrieval attribution are covered first, even though a team's most visible pain point in practice is usually further down the table, in screening, metrics, cost, or incident response.

The two disciplines feed each other constantly once an agent is live. A stored trace becomes the raw material an eval samples from — a curated slice of real runs is what a fixture set, a judge dataset, or a regression suite is actually built out of. An eval's verdict on a run — a low judge score, a failed step, an escalation to a human — is itself something worth tracing and monitoring over time, because a rising rate of any of those three is a signal in its own right, not just a per-run outcome to discard once scored. Treating tracing and evaluation as one continuous feedback loop, rather than two unrelated tools bolted onto an agent separately, is the organizing idea behind everything that follows.

Tracing spans for tool calls and retrieval steps

Tracing spans for tool calls and retrieval steps means giving each individual thing an agent does mid-run — call a model, fire off a tool, query a retrieval index — its own record, nested under whichever step triggered it, rather than logging everything into one undifferentiated stream. That nesting is what turns a bad final answer into a solvable problem: instead of re-reading an entire flat log line by line, a responder opens the one span that actually misbehaved and reads its inputs and outputs directly. OpenTelemetry's GenAI semantic conventions supply the vendor-neutral vocabulary most observability tooling now targets for this, and it remains a moving one: the specification's own governing group only started publishing it in April 2024, and as of July 2026 its attribute names still shipped under a Development-status flag rather than a finished release — worth remembering before treating any specific field name, from any source including this one, as permanent.

What the wire-level attribute names are actually called, and which fields belong on a well-instrumented span, is a deep-dive this pillar does not repeat here: see what should an AI agent's observability system capture for the field-by-field mechanics. At survey depth, the fact worth carrying forward is simpler and applies regardless of which attribute vocabulary a pipeline uses:

Tooling choice — a framework-agnostic OTel platform versus tracing built directly into an agent framework — sits downstream of getting these fields captured at all, and is covered in the sibling article linked above rather than repeated here.

One more property of the trace ID matters for everything downstream in this guide: because the same identifier threads through every child operation a run touches, including a handoff to a sub-agent, it is what lets a person or a script correlate activity across an entire multi-agent pipeline rather than treating each agent's output as an unrelated event stream. Judge screening, online metrics, cost telemetry, and postmortems below all assume this correlation already exists — none of them works cleanly bolted onto per-agent logging that never shares an ID across a handoff.

Should a production agent trace every run, or only a sample?

A production agent should trace every run at the tracing layer itself, because the marginal cost of writing a span is small and a run that was not traced cannot later be sampled for evaluation, escalated for review, or preserved for a postmortem no matter how important it turns out to be in hindsight. Sampling belongs one layer up, at the evaluation and screening stage: a live judge, a human reviewer, or a dashboard reads a deliberately chosen subset of the full trace stream, not the whole thing, because scoring every single run with a second model or a person does not scale and was never the point of screening in the first place. Conflating the two — under-tracing to save on logging cost, or over-screening every run with a judge to compensate for gaps in tracing — is a common way teams end up with neither reliable debugging data nor an affordable eval pipeline.

What does a retrieval-attribution log capture?

A retrieval-attribution log captures which specific documents or context chunks a retrieval step pulled into an agent's working context for a given run, so an operator can trace a wrong or fabricated answer back to what the agent actually saw rather than guessing whether the failure was a retrieval problem or a reasoning problem. No resource in this site's corpus names "retrieval-attribution logging" as a standalone practice; treat it here as the natural extension of the general span model to one specific operation type — agent-observability already lists a retrieval lookup as one of the operations a span records, alongside a model call, a tool call, and a sub-agent hop, and attribution logging is simply the discipline of keeping what that span retrieved, not just that it ran.

Two failure classes only a retrieval-attribution log can tell apart:

Distinguishing the two before assigning a fix is the entire value of logging retrieval spans specifically rather than folding them into a generic "tool call" span category that discards which content actually reached the model. For the deeper question of scoring retrieval quality itself — precision, recall, chunk relevance — see evaluating RAG systems in this site's separate RAG-in-production cluster; retrieval-attribution logging here is the observability precondition that makes that kind of scoring possible against live traffic, not a restatement of it.

Judge-based screening vs. human acceptance in production

Judge-based screening uses a second model to score or flag an agent's live output against a rubric, while human acceptance routes some share of those same outputs to a person for a manual accept-or-reject decision. In production, unlike in a pre-ship CI gate, both run continuously against real traffic rather than once against a fixed offline set. The judge model itself carries the same three documented failure modes regardless of whether it is screening a CI candidate or a live response: it favors certain answer positions in a comparison, it prefers a wordier, more formally worded answer over an equally correct but shorter one, and it rates its own model family's answers more generously than an equally correct answer from a different family. How to evaluate AI agents in CI covers that bias set in full for the pre-ship case; in production the same biases mean a live judge is best treated as a triage signal, not a final verdict, on anything a false negative would make expensive to miss.

Evaluating-ai-agents itself scopes human review narrowly: it names human judgment as the ground truth used to calibrate an automated evaluator, and as the right tool for scoring subjective, open-ended quality rather than for the routine bulk of live traffic — not as the default path for every response an agent produces, which cannot scale past a small fraction of production volume. One reasonable inference from that scoping, though no resource in this corpus documents a specific production screening architecture: route a judge's low-confidence or high-stakes flags to a human queue, and let confidently-scored output pass through unreviewed. That trade-off is worth stating honestly rather than glossing over — accepting it means some fraction of what passes through unreviewed is wrong in ways a human would have caught, and judge-based screening was never designed to eliminate that fraction, only to make the volume a small human team can actually cover.

Judge-based screening Human acceptance
Coverage Every sampled run, cheaply A small fraction of runs, deliberately chosen
Speed Near-instant, runs inline with production traffic Minutes to hours, depending on queue depth
Known bias risk Position, verbosity, and self-preference — inherited from the model doing the scoring None inherent, though a reviewer's own fatigue and inconsistency across a long queue are real, unstudied-here risks
Best suited to High-volume, lower-stakes traffic with a checkable or rubric-scoreable shape Low-volume, high-stakes, or genuinely open-ended traffic a rubric cannot fully capture

Neither column is a complete answer on its own — a judge alone inherits its documented biases at scale, and a human-only pipeline simply cannot cover production volume, which is why the two are almost always run together rather than as alternatives.

What online metrics should you track for a shipped AI agent?

Online metrics for a shipped AI agent are measurements pulled from live traffic rather than a fixed offline dataset: explicit user feedback, implicit behavioral signals, and the rate at which an interaction gets escalated away from the agent. They exist specifically to catch the distribution shift that an offline eval set, however well built, cannot see once real users start sending inputs its designers never anticipated. Evaluating-ai-agents itself draws exactly this line: online evaluation "catches distribution shift and contamination" that offline evaluation, run against a pre-computed reference set, structurally cannot.

No resource in this site's corpus defines a specific "escalation rate" metric or names a threshold for when a rising escalation share signals a problem. Treating it as worth tracking is a reasoned inference from that same online-versus-offline distinction — a rising share of interactions leaving an agent's control is exactly the kind of live-traffic signal the distinction says an offline set cannot catch — not a documented industry standard, and a team adopting it should set its own baseline rather than borrow a number from elsewhere.

Two other signal types fit the same online-eval framing:

Neither signal type is producible from a fixed offline dataset, because both only exist once a real user is on the other end of the interaction — which is exactly why evaluating-ai-agents treats online evaluation as slower and harder to reproduce than offline evaluation, but able to catch what offline evaluation cannot. None of these signals is actionable without the trace data from the sections above sitting underneath it: a rising escalation rate only becomes a fixable problem once it can be joined back to the specific spans, tool calls, and retrieved context that produced the escalated runs. For how reliable each of these signals actually is, see how to measure AI agent quality from live user feedback. For how to baseline those aggregates and tell a version change from a traffic shift, see how to detect quality drift in a production AI agent.

Cost telemetry: measuring what a production agent actually spends

Cost telemetry for a production AI agent means capturing what each span and each full trace actually cost in tokens and dollars as the agent runs — a measurement discipline, distinct from the separate question of how to bring that cost down. Agent-observability specifies per-span token usage and derived cost as one of the signals worth capturing on every span; agent-cost-latency-optimization states the reason plainly: a team "cannot optimize what" it does "not measure," which is why cost and latency belong on every span before any lever gets pulled, not after.

That resource's four-tier lever taxonomy — token-level, request-level, model-level, and architecture-level fixes — is the separate discipline of reducing what an agent costs once telemetry has told you where the spend actually goes; see agent cost and latency optimization for those levers in full, since restating them here would blur a measurement question into an optimization one. Practical cost telemetry aggregates the same per-span number along at least three axes an individual span cannot show on its own, though no resource prescribes this specific breakdown — it follows directly from the per-span field agent-observability already specifies, aggregated along the axes an operator actually needs to act on:

Why measure at the trace level rather than trusting an average per-call price: agent-cost-latency-optimization documents fan-out — a pipeline that dispatches several sub-agents, each making many calls — as the single fastest way total spend departs from what any one call's price would suggest, and recommends measuring steps-per-task before optimizing individual calls at all. Per-call cost telemetry alone cannot see fan-out; only a trace-level rollup, summing every span under one run, shows a five-sub-agent pipeline actually costing fifty times a single-agent baseline rather than five times it.

How much trace and eval data should you keep?

How much trace and eval data to keep is a question no resource in this site's corpus answers with a specific retention number, and the honest position is that the right window depends on a deployment's own incident-investigation needs, storage budget, and any regulatory retention requirement it carries — not a single default every team should copy. Two things hold regardless of the specific number a team lands on. First, retaining enough of a trace to reconstruct a full run outweighs compressing or truncating early, because a shortened trace forecloses exactly the incident-investigation and fixture-building work the sections above depend on — the gap is invisible until the one time it matters. Second, sampling for judge screening, dashboards, and eval-fixture curation needs a much smaller, deliberately chosen subset than raw trace retention needs, since none of those consumers reads the full stream — conflating "how much to store" with "how much to actively screen" leads teams to either over-spend on judge calls against traffic that never needed scoring, or under-store the raw traces those same judges and postmortems need to work from later.

How do incident postmortems become evaluation fixtures?

An incident postmortem becomes an evaluation fixture when a team takes the specific input or trajectory that caused a live failure and adds it, as a new labeled test case, to the same suite that gates every future release — turning one incident into a permanent check rather than a story nobody re-reads. No resource in this site's corpus documents a specific postmortem process's cadence, retention window, or numeric threshold for promoting a finding into a fixture; the survey below covers the mechanism, not a prescribed procedure, following the same operational-layer treatment how to build an incident response runbook for AI agent failures — in the completed agent-reliability cluster — already applies to the failure classes that produce these incidents in the first place.

The mechanism leans on the same trace-freezing habit incident response needs for an entirely separate reason: exporting the evidence before a recovery step gets a chance to wipe it doubles as preserving the raw material a fixture is built from, so a team that never learned to freeze a trace for its own investigation has nothing to build a fixture out of later either. This is not a new idea invented for incident response specifically — testing-ai-agents itself notes that traces from production runs can seed new cassettes and test cases for a CI suite; a postmortem fixture is simply the deliberate, incident-triggered version of that same general practice. Once a trace is preserved, converting it into a fixture is the same operation as building any other eval case: pair the recorded input and trajectory with the correct expected outcome, and run it through whichever layer of the eval suite would have caught it — how to evaluate AI agents in CI covers what that suite actually looks like, and this section points there rather than re-deriving it.

Closing the loop: how tracing, evaluation, and incident response reinforce each other

Closing the loop between tracing, evaluation, and incident response means treating one stored trace as the single artifact all three disciplines share, rather than building separate ad hoc logging for each. A trace captures what happened; a judge or a human screens a sample of what happened and produces a verdict; an online metric aggregates those verdicts, plus feedback and escalation signals, into a trend a team actually watches; a trend crossing an unacceptable line becomes an incident; an incident's frozen trace becomes a postmortem; and a postmortem's output becomes a new fixture that feeds straight back into the CI suite covered in full by how to evaluate AI agents in CI — which is exactly what closes the loop, since the next release is now tested against a failure the previous one never anticipated.

None of the six disciplines above is exotic engineering on its own — trace IDs, judge models, feedback widgets, cost dashboards, and postmortems are all established practice well outside agentic systems. What is specific to an AI agent is that a single bad run can pass through every one of them in sequence, and a team that only builds three of the six has a loop with a gap exactly where the next real incident will find it.

A production observability-and-evaluation checklist

Before calling an agent's post-ship observability and evaluation surface complete, it should clear this list:

The full agent observability and evaluation cluster

Eleven sub-articles make up the agent-observability-evaluation cluster, each going deeper on one piece of this guide, and the cluster is complete as of October 2026:

Sources and further reading

Every factual claim above is carried, with its primary source, by a reference resource in this site's corpus:

For the two disciplines this guide deliberately stays at survey depth on, see how to evaluate AI agents in CI and what should an AI agent's observability system capture. Both live in the completed agent-reliability cluster; this guide sits next to it, not on top of it. Agents: this guide has a Markdown variant at /articles/agent-observability-and-evaluation.md, and the whole editorial layer is indexed as JSON at /api/articles.json.

Frequently asked questions

What is the difference between AI agent observability and AI agent evaluation?
AI agent observability is the practice of capturing what an agent actually did on a given run — its trace of model calls, tool calls, and retrieval steps — while AI agent evaluation is the separate practice of judging whether what it did was correct, safe, or good enough; the two disciplines run on the same underlying trace data but ask different questions of it, and a production system needs both rather than treating either as a substitute for the other.
What is LLM tracing, and is it the same as AI agent observability?
LLM tracing is the narrower practice of recording individual model calls — prompts, completions, token counts, and latency — as they happen, while AI agent observability extends the same idea across a full multi-step run: a trace nests LLM calls alongside tool calls, retrieval steps, and sub-agent hops under one shared run ID, so LLM tracing on its own captures only one of several operation types an agent trace needs in order to reconstruct what actually happened.
Does judge-based screening replace human review for a production AI agent?
No — a judge model carries the same position, verbosity, and self-preference biases in production that are documented for offline evaluation, so treating every judge verdict as final rather than routing its low-confidence or high-stakes flags to a human queue accepts that some share of unreviewed output is wrong in ways a person would have caught; judge-based screening scales what a small human review team can cover, but it does not eliminate the need for that team.
Should cost telemetry be used to optimize an AI agent's spend?
Cost telemetry itself is a measurement discipline, not an optimization one: capturing what each span and trace actually cost in tokens and dollars is the precondition for optimization, but reducing that cost through token-level, request-level, model-level, or architecture-level levers is a separate, later step documented in the dedicated agent cost and latency optimization reference rather than in the telemetry itself.
How does an AI agent incident postmortem turn into an evaluation fixture?
An incident postmortem turns into an evaluation fixture when a team takes the specific input or trajectory that caused a live failure, pairs it with the correct expected outcome, and adds it as a new test case to the same suite that gates every future release, which requires the trace to have been preserved before any rollback or restart could overwrite the evidence a fixture needs.

#agents #observability #evaluation #tracing #production #monitoring #cost

Put this corpus inside your own agents

This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.

Agents: this page as Markdown · JSON · offers at /api/pricing.json · payment methods at /api/payment.json · single-resource access via HTTP 402 (how that works)