AI Agent Observability and the Production Evaluation Playbook
AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.
- AI agent observability and AI agent evaluation are two distinct disciplines built on the same trace data: observability captures what an agent actually did on a given run, and evaluation judges whether what it did was any good.
- A retrieval-attribution log records which specific documents or context chunks a retrieval step pulled into an agent's working context, letting an operator tell a retrieval failure apart from a generation failure instead of guessing between the two.
- A live judge model screening production AI agent output carries the same position, verbosity, and self-preference biases documented for offline evaluation, which is why a production system typically routes its low-confidence or high-stakes flags to a human queue rather than trusting every verdict unreviewed.
- Cost telemetry for a production AI agent means capturing what each span and trace actually spent in tokens and dollars, a measurement discipline distinct from the separate question of how to reduce that spend.
- An incident postmortem for an AI agent becomes durable only once the specific input or trajectory that caused the failure is added as a new fixture to the same evaluation suite that gates future releases, turning one incident into a permanent check.
- No resource in this site's corpus defines a standard "escalation rate" metric for AI agents as of September 2026, so a team tracking one should set its own baseline rather than borrow a number from elsewhere.
An agent that fails loudly on day one is a cheaper problem than one that quietly does the wrong thing for three days before anyone notices — observability is what lets an operator tell the difference, and evaluation is what tells them whether the difference matters. This guide surveys both disciplines as they operate on an AI agent that is already live: tracing spans for tool calls and retrieval steps, screening live output with a judge model versus routing it to a human, the online metrics worth watching once real traffic replaces a benchmark set, what a running agent actually costs, and how an incident turns into a fixture the next release has to pass.
Two things this guide deliberately does not do. It does not re-derive pre-ship, CI-gating evaluation methodology — the public benchmark landscape, pass@k versus pass^k, the three-tier CI test pyramid — which how to evaluate AI agents in CI already covers in full; this guide treats evaluation only at the depth needed to place it next to observability, and points there for anything that gates a release before it ships. It also does not re-derive trace and span field mechanics — the specific gen_ai.* OpenTelemetry attribute names, or the five-field-per-span checklist — which what should an AI agent's observability system capture already covers in full; this guide treats tracing only at the depth needed to make the topic legible, and points there for the wire-level detail.
The rest of this guide moves through six disciplines in roughly the order a live agent actually produces the data they need: tracing, retrieval attribution, judge-versus-human screening, online metrics, cost telemetry, and postmortem-to-fixture feedback.
What is AI agent observability, and how does it differ from evaluation?
AI agent observability is the practice of capturing what an agent actually did on a given run — its trace of model calls, tool calls, and retrievals — while AI agent evaluation is the separate practice of judging whether what it did was correct, safe, or good enough. Observability answers "what happened": a trace ties one full run under a single ID, and each span inside it nests a model call, tool call, retrieval step, or sub-agent hop, as agent observability and tracing documents in full. Evaluation answers "was it any good": scoring a run, or a batch of runs, against a rubric, a benchmark, or a reference answer, as evaluating AI agents documents in full.
The table below is a map of the six disciplines this guide covers and how they relate, not a substitute for reading each section — it exists so a reader can place a specific problem (a cost spike, a bad live answer, a repeat incident) against the discipline that actually owns it before diving into the detail:
| Discipline | Question it answers | What feeds it |
|---|---|---|
| Tracing | What did the agent actually do on this one run? | Nested spans under a shared trace ID |
| Retrieval attribution | What context produced this specific output? | Retrieval-step spans, kept distinct from other tool spans |
| Judge/human screening | Was this specific live output any good? | A sampled slice of live traffic plus a rubric |
| Online metrics | Is behavior drifting across many runs, not just one? | Feedback, behavioral signals, escalation share |
| Cost telemetry | What did this run or trace actually spend? | Per-span token counts converted to cost |
| Incident postmortem | What should never happen the same way twice? | A frozen trace from a run that already failed |
Every row after the first two depends on the trace data the first two produce — which is why tracing and retrieval attribution are covered first, even though a team's most visible pain point in practice is usually further down the table, in screening, metrics, cost, or incident response.
The two disciplines feed each other constantly once an agent is live. A stored trace becomes the raw material an eval samples from — a curated slice of real runs is what a fixture set, a judge dataset, or a regression suite is actually built out of. An eval's verdict on a run — a low judge score, a failed step, an escalation to a human — is itself something worth tracing and monitoring over time, because a rising rate of any of those three is a signal in its own right, not just a per-run outcome to discard once scored. Treating tracing and evaluation as one continuous feedback loop, rather than two unrelated tools bolted onto an agent separately, is the organizing idea behind everything that follows.
Tracing spans for tool calls and retrieval steps
Tracing spans for tool calls and retrieval steps means giving each individual thing an agent does mid-run — call a model, fire off a tool, query a retrieval index — its own record, nested under whichever step triggered it, rather than logging everything into one undifferentiated stream. That nesting is what turns a bad final answer into a solvable problem: instead of re-reading an entire flat log line by line, a responder opens the one span that actually misbehaved and reads its inputs and outputs directly. OpenTelemetry's GenAI semantic conventions supply the vendor-neutral vocabulary most observability tooling now targets for this, and it remains a moving one: the specification's own governing group only started publishing it in April 2024, and as of July 2026 its attribute names still shipped under a Development-status flag rather than a finished release — worth remembering before treating any specific field name, from any source including this one, as permanent.
What the wire-level attribute names are actually called, and which fields belong on a well-instrumented span, is a deep-dive this pillar does not repeat here: see what should an AI agent's observability system capture for the field-by-field mechanics. At survey depth, the fact worth carrying forward is simpler and applies regardless of which attribute vocabulary a pipeline uses:
- A tool-call span is only useful for either observability or evaluation if it records the call's inputs, outputs, and outcome independently of what the model's own final answer claims happened.
- An agent that reports success on a tool call that actually errored is a failure mode that neither a flat log nor an untraced pipeline will ever surface — only a span attached to that specific call, with its own recorded outcome, catches it.
- The same nesting that makes a trace legible to a human debugger is what makes it legible to an automated eval later: a fixture built from an incomplete or flattened trace can only test what the trace happened to preserve.
Tooling choice — a framework-agnostic OTel platform versus tracing built directly into an agent framework — sits downstream of getting these fields captured at all, and is covered in the sibling article linked above rather than repeated here.
One more property of the trace ID matters for everything downstream in this guide: because the same identifier threads through every child operation a run touches, including a handoff to a sub-agent, it is what lets a person or a script correlate activity across an entire multi-agent pipeline rather than treating each agent's output as an unrelated event stream. Judge screening, online metrics, cost telemetry, and postmortems below all assume this correlation already exists — none of them works cleanly bolted onto per-agent logging that never shares an ID across a handoff.
Should a production agent trace every run, or only a sample?
A production agent should trace every run at the tracing layer itself, because the marginal cost of writing a span is small and a run that was not traced cannot later be sampled for evaluation, escalated for review, or preserved for a postmortem no matter how important it turns out to be in hindsight. Sampling belongs one layer up, at the evaluation and screening stage: a live judge, a human reviewer, or a dashboard reads a deliberately chosen subset of the full trace stream, not the whole thing, because scoring every single run with a second model or a person does not scale and was never the point of screening in the first place. Conflating the two — under-tracing to save on logging cost, or over-screening every run with a judge to compensate for gaps in tracing — is a common way teams end up with neither reliable debugging data nor an affordable eval pipeline.
What does a retrieval-attribution log capture?
A retrieval-attribution log captures which specific documents or context chunks a retrieval step pulled into an agent's working context for a given run, so an operator can trace a wrong or fabricated answer back to what the agent actually saw rather than guessing whether the failure was a retrieval problem or a reasoning problem. No resource in this site's corpus names "retrieval-attribution logging" as a standalone practice; treat it here as the natural extension of the general span model to one specific operation type — agent-observability already lists a retrieval lookup as one of the operations a span records, alongside a model call, a tool call, and a sub-agent hop, and attribution logging is simply the discipline of keeping what that span retrieved, not just that it ran.
Two failure classes only a retrieval-attribution log can tell apart:
- A retrieval step that returned no relevant context at all — this is an index or query problem, and the fix lives in the retrieval pipeline itself.
- A retrieval step that returned relevant context the model then ignored or contradicted — this is a generation problem, and no amount of re-indexing will fix it.
Distinguishing the two before assigning a fix is the entire value of logging retrieval spans specifically rather than folding them into a generic "tool call" span category that discards which content actually reached the model. For the deeper question of scoring retrieval quality itself — precision, recall, chunk relevance — see evaluating RAG systems in this site's separate RAG-in-production cluster; retrieval-attribution logging here is the observability precondition that makes that kind of scoring possible against live traffic, not a restatement of it.
Judge-based screening vs. human acceptance in production
Judge-based screening uses a second model to score or flag an agent's live output against a rubric, while human acceptance routes some share of those same outputs to a person for a manual accept-or-reject decision. In production, unlike in a pre-ship CI gate, both run continuously against real traffic rather than once against a fixed offline set. The judge model itself carries the same three documented failure modes regardless of whether it is screening a CI candidate or a live response: it favors certain answer positions in a comparison, it prefers a wordier, more formally worded answer over an equally correct but shorter one, and it rates its own model family's answers more generously than an equally correct answer from a different family. How to evaluate AI agents in CI covers that bias set in full for the pre-ship case; in production the same biases mean a live judge is best treated as a triage signal, not a final verdict, on anything a false negative would make expensive to miss.
Evaluating-ai-agents itself scopes human review narrowly: it names human judgment as the ground truth used to calibrate an automated evaluator, and as the right tool for scoring subjective, open-ended quality rather than for the routine bulk of live traffic — not as the default path for every response an agent produces, which cannot scale past a small fraction of production volume. One reasonable inference from that scoping, though no resource in this corpus documents a specific production screening architecture: route a judge's low-confidence or high-stakes flags to a human queue, and let confidently-scored output pass through unreviewed. That trade-off is worth stating honestly rather than glossing over — accepting it means some fraction of what passes through unreviewed is wrong in ways a human would have caught, and judge-based screening was never designed to eliminate that fraction, only to make the volume a small human team can actually cover.
| Judge-based screening | Human acceptance | |
|---|---|---|
| Coverage | Every sampled run, cheaply | A small fraction of runs, deliberately chosen |
| Speed | Near-instant, runs inline with production traffic | Minutes to hours, depending on queue depth |
| Known bias risk | Position, verbosity, and self-preference — inherited from the model doing the scoring | None inherent, though a reviewer's own fatigue and inconsistency across a long queue are real, unstudied-here risks |
| Best suited to | High-volume, lower-stakes traffic with a checkable or rubric-scoreable shape | Low-volume, high-stakes, or genuinely open-ended traffic a rubric cannot fully capture |
Neither column is a complete answer on its own — a judge alone inherits its documented biases at scale, and a human-only pipeline simply cannot cover production volume, which is why the two are almost always run together rather than as alternatives.
What online metrics should you track for a shipped AI agent?
Online metrics for a shipped AI agent are measurements pulled from live traffic rather than a fixed offline dataset: explicit user feedback, implicit behavioral signals, and the rate at which an interaction gets escalated away from the agent. They exist specifically to catch the distribution shift that an offline eval set, however well built, cannot see once real users start sending inputs its designers never anticipated. Evaluating-ai-agents itself draws exactly this line: online evaluation "catches distribution shift and contamination" that offline evaluation, run against a pre-computed reference set, structurally cannot.
No resource in this site's corpus defines a specific "escalation rate" metric or names a threshold for when a rising escalation share signals a problem. Treating it as worth tracking is a reasoned inference from that same online-versus-offline distinction — a rising share of interactions leaving an agent's control is exactly the kind of live-traffic signal the distinction says an offline set cannot catch — not a documented industry standard, and a team adopting it should set its own baseline rather than borrow a number from elsewhere.
Two other signal types fit the same online-eval framing:
- Explicit feedback — a thumbs-up/down, a follow-up correction, a user re-asking the same question in different words.
- Implicit behavioral signals — session abandonment, an immediate retry of the same task, a support ticket filed shortly after an agent interaction.
Neither signal type is producible from a fixed offline dataset, because both only exist once a real user is on the other end of the interaction — which is exactly why evaluating-ai-agents treats online evaluation as slower and harder to reproduce than offline evaluation, but able to catch what offline evaluation cannot. None of these signals is actionable without the trace data from the sections above sitting underneath it: a rising escalation rate only becomes a fixable problem once it can be joined back to the specific spans, tool calls, and retrieved context that produced the escalated runs. For how reliable each of these signals actually is, see how to measure AI agent quality from live user feedback. For how to baseline those aggregates and tell a version change from a traffic shift, see how to detect quality drift in a production AI agent.
Cost telemetry: measuring what a production agent actually spends
Cost telemetry for a production AI agent means capturing what each span and each full trace actually cost in tokens and dollars as the agent runs — a measurement discipline, distinct from the separate question of how to bring that cost down. Agent-observability specifies per-span token usage and derived cost as one of the signals worth capturing on every span; agent-cost-latency-optimization states the reason plainly: a team "cannot optimize what" it does "not measure," which is why cost and latency belong on every span before any lever gets pulled, not after.
That resource's four-tier lever taxonomy — token-level, request-level, model-level, and architecture-level fixes — is the separate discipline of reducing what an agent costs once telemetry has told you where the spend actually goes; see agent cost and latency optimization for those levers in full, since restating them here would blur a measurement question into an optimization one. Practical cost telemetry aggregates the same per-span number along at least three axes an individual span cannot show on its own, though no resource prescribes this specific breakdown — it follows directly from the per-span field agent-observability already specifies, aggregated along the axes an operator actually needs to act on:
- Cost per task type, so a team can see which categories of request are expensive by nature rather than by regression.
- Cost per tenant or user segment, so one unusually heavy caller does not get averaged away into an overall number that looks fine.
- Cost trend over time at a fixed task mix, so a genuine regression is distinguishable from ordinary demand growth.
Why measure at the trace level rather than trusting an average per-call price: agent-cost-latency-optimization documents fan-out — a pipeline that dispatches several sub-agents, each making many calls — as the single fastest way total spend departs from what any one call's price would suggest, and recommends measuring steps-per-task before optimizing individual calls at all. Per-call cost telemetry alone cannot see fan-out; only a trace-level rollup, summing every span under one run, shows a five-sub-agent pipeline actually costing fifty times a single-agent baseline rather than five times it.
How much trace and eval data should you keep?
How much trace and eval data to keep is a question no resource in this site's corpus answers with a specific retention number, and the honest position is that the right window depends on a deployment's own incident-investigation needs, storage budget, and any regulatory retention requirement it carries — not a single default every team should copy. Two things hold regardless of the specific number a team lands on. First, retaining enough of a trace to reconstruct a full run outweighs compressing or truncating early, because a shortened trace forecloses exactly the incident-investigation and fixture-building work the sections above depend on — the gap is invisible until the one time it matters. Second, sampling for judge screening, dashboards, and eval-fixture curation needs a much smaller, deliberately chosen subset than raw trace retention needs, since none of those consumers reads the full stream — conflating "how much to store" with "how much to actively screen" leads teams to either over-spend on judge calls against traffic that never needed scoring, or under-store the raw traces those same judges and postmortems need to work from later.
How do incident postmortems become evaluation fixtures?
An incident postmortem becomes an evaluation fixture when a team takes the specific input or trajectory that caused a live failure and adds it, as a new labeled test case, to the same suite that gates every future release — turning one incident into a permanent check rather than a story nobody re-reads. No resource in this site's corpus documents a specific postmortem process's cadence, retention window, or numeric threshold for promoting a finding into a fixture; the survey below covers the mechanism, not a prescribed procedure, following the same operational-layer treatment how to build an incident response runbook for AI agent failures — in the completed agent-reliability cluster — already applies to the failure classes that produce these incidents in the first place.
The mechanism leans on the same trace-freezing habit incident response needs for an entirely separate reason: exporting the evidence before a recovery step gets a chance to wipe it doubles as preserving the raw material a fixture is built from, so a team that never learned to freeze a trace for its own investigation has nothing to build a fixture out of later either. This is not a new idea invented for incident response specifically — testing-ai-agents itself notes that traces from production runs can seed new cassettes and test cases for a CI suite; a postmortem fixture is simply the deliberate, incident-triggered version of that same general practice. Once a trace is preserved, converting it into a fixture is the same operation as building any other eval case: pair the recorded input and trajectory with the correct expected outcome, and run it through whichever layer of the eval suite would have caught it — how to evaluate AI agents in CI covers what that suite actually looks like, and this section points there rather than re-deriving it.
Closing the loop: how tracing, evaluation, and incident response reinforce each other
Closing the loop between tracing, evaluation, and incident response means treating one stored trace as the single artifact all three disciplines share, rather than building separate ad hoc logging for each. A trace captures what happened; a judge or a human screens a sample of what happened and produces a verdict; an online metric aggregates those verdicts, plus feedback and escalation signals, into a trend a team actually watches; a trend crossing an unacceptable line becomes an incident; an incident's frozen trace becomes a postmortem; and a postmortem's output becomes a new fixture that feeds straight back into the CI suite covered in full by how to evaluate AI agents in CI — which is exactly what closes the loop, since the next release is now tested against a failure the previous one never anticipated.
None of the six disciplines above is exotic engineering on its own — trace IDs, judge models, feedback widgets, cost dashboards, and postmortems are all established practice well outside agentic systems. What is specific to an AI agent is that a single bad run can pass through every one of them in sequence, and a team that only builds three of the six has a loop with a gap exactly where the next real incident will find it.
A production observability-and-evaluation checklist
Before calling an agent's post-ship observability and evaluation surface complete, it should clear this list:
- Every model call, tool call, and retrieval step in a run is captured as its own record, nested beneath whichever step triggered it, all sharing one trace ID for the run.
- Retrieval spans record what was actually retrieved, not just that a retrieval happened, so a wrong answer can be traced to a retrieval failure or a generation failure specifically.
- A live judge screens production output against a documented rubric, with its low-confidence or high-stakes flags routed to a human queue rather than trusted unreviewed.
- Online metrics — explicit feedback, implicit behavioral signals, and an escalation share — are tracked against a team's own baseline, not a borrowed industry number.
- Cost and latency are captured per span and per trace, and aggregated by task type, tenant, and trend, separately from any optimization work done with that same data.
- Every incident gets its trace exported before any recovery step — a rollback, a restart — has a chance to overwrite it, and every postmortem ships at least one new fixture into the CI evaluation suite.
The full agent observability and evaluation cluster
Eleven sub-articles make up the agent-observability-evaluation cluster, each going deeper on one piece of this guide, and the cluster is complete as of October 2026:
- Distributed tracing for multi-agent AI systems extends the tracing section to handoffs between agents and the silent-fork failure.
- Designing an LLM-as-judge pipeline for production AI agents turns the judge-versus-human section into a confidence-routed review design.
- How to track AI agent costs in production expands the cost-telemetry section into price tables, attribution and reconciliation.
- How to turn an AI agent incident into an evaluation test case walks the postmortem-to-fixture procedure step by step.
- How to log RAG retrieval in production for debugging agent answers details what a retrieval-attribution log records and why.
- How to measure AI agent quality from live user feedback covers explicit and implicit feedback signals and their biases.
- How to detect quality drift in a production AI agent turns the online-metrics section into denominators, baselines and a drift triage order.
- How to sample and retain production AI agent traces expands the keep-or-drop question into sampling tiers and retention.
- How to keep trace data when an AI agent crashes mid-run covers run-start records and orphan detection for runs that die before export.
- How to choose an LLM observability platform for AI agents gives the decision criteria for picking a tracing backend.
- How to redact PII from AI agent traces covers redactor placement, verification and cleanup of leaked spans.
Sources and further reading
Every factual claim above is carried, with its primary source, by a reference resource in this site's corpus:
- Trace/span structure, OpenTelemetry GenAI conventions, and the tooling landscape: /resources/agent-observability
- Agent eval methodology, LLM-as-judge biases, and the online-vs-offline distinction: /resources/evaluating-ai-agents
- CI test-pyramid mechanics, plus the note that production traces can seed new test cases: /resources/testing-ai-agents
- Cost and latency measurement plus the separate four-tier optimization-lever taxonomy: /resources/agent-cost-latency-optimization
For the two disciplines this guide deliberately stays at survey depth on, see how to evaluate AI agents in CI and what should an AI agent's observability system capture. Both live in the completed agent-reliability cluster; this guide sits next to it, not on top of it. Agents: this guide has a Markdown variant at /articles/agent-observability-and-evaluation.md, and the whole editorial layer is indexed as JSON at /api/articles.json.
Frequently asked questions
- What is the difference between AI agent observability and AI agent evaluation?
- AI agent observability is the practice of capturing what an agent actually did on a given run — its trace of model calls, tool calls, and retrieval steps — while AI agent evaluation is the separate practice of judging whether what it did was correct, safe, or good enough; the two disciplines run on the same underlying trace data but ask different questions of it, and a production system needs both rather than treating either as a substitute for the other.
- What is LLM tracing, and is it the same as AI agent observability?
- LLM tracing is the narrower practice of recording individual model calls — prompts, completions, token counts, and latency — as they happen, while AI agent observability extends the same idea across a full multi-step run: a trace nests LLM calls alongside tool calls, retrieval steps, and sub-agent hops under one shared run ID, so LLM tracing on its own captures only one of several operation types an agent trace needs in order to reconstruct what actually happened.
- Does judge-based screening replace human review for a production AI agent?
- No — a judge model carries the same position, verbosity, and self-preference biases in production that are documented for offline evaluation, so treating every judge verdict as final rather than routing its low-confidence or high-stakes flags to a human queue accepts that some share of unreviewed output is wrong in ways a person would have caught; judge-based screening scales what a small human review team can cover, but it does not eliminate the need for that team.
- Should cost telemetry be used to optimize an AI agent's spend?
- Cost telemetry itself is a measurement discipline, not an optimization one: capturing what each span and trace actually cost in tokens and dollars is the precondition for optimization, but reducing that cost through token-level, request-level, model-level, or architecture-level levers is a separate, later step documented in the dedicated agent cost and latency optimization reference rather than in the telemetry itself.
- How does an AI agent incident postmortem turn into an evaluation fixture?
- An incident postmortem turns into an evaluation fixture when a team takes the specific input or trajectory that caused a live failure, pairs it with the correct expected outcome, and adds it as a new test case to the same suite that gates every future release, which requires the trace to have been preserved before any rollback or restart could overwrite the evidence a fixture needs.
This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.