ChangeGamer

← All guides · Agent observability and evaluation

How to Sample and Retain Production AI Agent Traces

Part 8 of Agent observability and evaluation · 1,386 words · ~6 min read · published 2026-10-04 · updated 2026-10-04 · Markdown variant

How to sample and retain production AI agent traces: head vs tail sampling, keep-all-errors plus a random baseline, whole-trace decisions for multi-agent runs, retention tiers and redaction before the clock starts.

In short

  • Production agent traces should be sampled with a tail-based rule that keeps every errored, flagged or slow run plus a random baseline slice, because a head-based coin flip discards the runs you most need.
  • A tail-based keep decision needs the whole run buffered until it ends, so the cost of tail sampling is memory and delayed export rather than lost data.
  • A multi-agent trace should be kept or dropped as one unit, since a trace missing the sub-agent that failed cannot explain the failure.
  • Redaction of personal data must happen before a trace is stored, because the retention clock then starts on already-sanitized data and deletion requests have less to find.
  • No resource in the reference corpus publishes a trace sampling rate or retention window as of October 2026, so each team must set both from its own incident and storage needs.

Part of the AI Agent Observability and the Production Evaluation Playbook guide.


Production agent traces should be sampled at the point where the run ends, keeping every errored, flagged or slow run and a random baseline slice, and retained in tiers that start their clock after redaction. The pillar says to trace every run and screen a sample. This article covers the storage side of that advice. It is reasoned design as of October 2026. No resource in the reference corpus publishes a sampling rate or retention window, and none appears below. Search demand is validated for the broader topic of AI agent observability, not for this wording.

Should you sample before the run or after it?

Sample after the run for anything you want to debug, because only a decision made at run end can see whether the run failed. The two approaches differ in when the keep decision happens:

Head-based Tail-based
Decision time When the trace starts When the run finishes
Knows the outcome No Yes
Can keep all errors Only by chance Yes, by rule
Buffering needed None The full run until the decision
Typical use Cheap background volume Debugging and evaluation sets

A head-based coin flip is statistically fair and operationally blind. If errors are rare, most of them are dropped along with the healthy runs. Tail-based sampling pays for its accuracy in memory, because every in-flight run is held until it ends. Whether your tracing stack implements either one, and with what limits, is tool-specific. This article does not claim sampler behavior for any named tool, so check your own stack's documentation.

What should a keep rule always retain?

A keep rule should always retain errored, flagged and slow runs, plus a random baseline slice of everything else. The first three are the runs you will open during an incident. The baseline is what lets you compare them against normal.

Without the baseline, your stored set contains only failures, and every conclusion drawn from it is biased. You cannot tell what a bad run does differently from a good one if no good runs were kept. The baseline slice also feeds evaluation: a stored sample is the raw material for a dataset, as the agent-observability resource notes. The rate of the baseline slice is not published anywhere in the corpus, so pick it from your volume and write it down beside the rule.

How do you handle a tail decision for a long run?

Handle a tail decision by buffering spans in the process or collector until the root span closes, then applying the keep rule to the finished run. Agent runs are long and branching, so the buffer is the practical cost to plan for:

  1. Bound the buffer. Set a maximum run duration or span count after which the run is force-decided, so a stuck run cannot hold memory forever.
  2. Decide on the root. Evaluate error, flag and duration rules once the outermost span ends.
  3. Export or discard whole. Never send half a run.
  4. Record the rule that fired. Store which rule kept the trace, so later analysis can correct for the fact that errors are over-represented.

If the process dies mid-run, the buffered spans die with it. For crashes, you may want a separate lightweight record emitted at run start, which is a design choice rather than anything the corpus specifies.

How should a multi-agent trace be kept or dropped?

A multi-agent trace should be kept or dropped as a single unit, using the outcome of the whole tree. Sampling each agent's spans independently produces fragments in which the delegating agent is present and the failing sub-agent is gone, or the reverse.

The mechanism that makes this possible is a shared trace ID across every handoff, which the agent-observability resource describes as what enables cross-agent debugging. Passing that ID, and the keep decision, across process boundaries is covered in multi-agent trace propagation. The rule of thumb is simple: if any agent in the tree errors or is flagged, the whole tree is kept.

What does cost versus debuggability look like?

Cost versus debuggability comes down to which fields you pay to store, not only how many runs. A run's metadata, such as ID, duration, outcome and versions, is small. Its payloads, such as full tool inputs and outputs, are large and sensitive.

One reasonable split is to keep metadata for every run and payloads only for runs the keep rule selects. That keeps counts honest without storing everything. Spend is the exception to sampling: tokens and cost must be counted from every run, which is covered in cost telemetry. A sampled cost number understates exactly the expensive runs you need to see.

How should retention tiers be shaped?

Retention should be tiered by how sensitive and how large each class of data is, with a shorter life for raw payloads than for metadata. This is design shape only. The corpus gives no numbers and neither does this article.

Tier Contents Relative life
Metadata Run ID, versions, outcome, duration, counts Longest
Redacted trace Full span tree with sanitized payloads Medium
Raw payload Unredacted bodies, if you keep any at all Shortest, or none

The pillar already says retaining enough to reconstruct a run outweighs truncating early. The tiers are how you honor that without keeping sensitive bodies forever. Kept failures can become permanent regression checks through the incident-to-fixture loop, which is a reason to keep flagged runs in the redacted tier longer than baseline runs.

For the dashboards that read this data, per-tool views are covered in MCP server observability, and aggregate quality monitoring in online agent metrics. Field-level capture is in observability for reliability.

Why redact before the retention clock starts?

Redact before storage, because a retention window applied to unredacted traces is a window during which personal data sits in your trace store. The data-privacy-for-agents resource says to redact personal data before writing observability data, and to apply redaction at the exporter layer so every downstream sink receives only sanitized data. The agent-observability resource adds that prompt and completion bodies are off by default for PII safety.

Provider retention shows why this matters even outside your own store. The data-privacy-for-agents resource states that OpenAI keeps standard API logs up to 30 days, and that Anthropic reduced its standard retention from 30 to 7 days as of September 2025. Those are provider-side logs, and they do not shorten the life of traces in your store. Your own window is separate, and it should start from redacted data.

What does ChangeGamer do that resembles a retention window?

ChangeGamer sets expiry at write time rather than running a cleanup job, although it runs no agent trace store and this is an analogy only. In changegamer/worker/index.ts (lines 534-538), the session-to-key mapping is written to KV with { expirationTtl: 2592000 }, which is 30 days, and the corresponding resource documents that the mapping is retained for 30 days. The store enforces the window, so no process has to remember to delete anything.

The transferable idea is narrow: attach the lifetime when the record is written. For traces, that means the tier is chosen by the keep rule at export time, which is the same moment redaction runs. Nothing here is a measurement of an agent system.

What the corpus does not tell you

The corpus gives no sampling rate, buffer size, slow-run limit or retention window for agent traces, as of October 2026, and this article has not supplied any. The structure, a tail decision, a whole-trace unit, a baseline slice and redaction before retention, is reasoned from the resources cited and general practice. Treat it as a hypothesis, and revisit it after the first incident in which a trace you needed was not there.

Frequently asked questions

What is the difference between head-based and tail-based sampling for AI agent traces?
Head-based sampling decides whether to keep a trace when the run starts, before anything is known about it. Tail-based sampling decides after the run ends, so it can keep every error, flagged or slow run. The trade-off is that tail sampling must buffer the whole run until the decision.
How long should I keep production AI agent traces?
No resource in the reference corpus publishes a trace retention window as of October 2026, so the number is yours to set. Choose it from how long incidents take to surface, your storage budget and any regulatory retention rule, and keep a shorter window for raw payloads than for metadata.
Should I sample agent traces to save money?
Sampling agent traces to save storage is reasonable, but sampling cost telemetry is not, because spend must be counted from every run. Keep the full run record for cost and counts, and sample only the heavy trace payloads, always retaining errors and flagged runs.
Can I store raw prompts in agent traces?
Raw prompts in agent traces should be avoided by default. The agent-observability resource says prompt and completion bodies are off by default for PII safety, and data-privacy-for-agents says to redact personal data before writing observability data, at the exporter layer.

#agents #observability #tracing #sampling #retention #privacy

Put this corpus inside your own agents

This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.

Agents: this page as Markdown · JSON · offers at /api/pricing.json · payment methods at /api/payment.json · single-resource access via HTTP 402 (how that works)