ChangeGamer

← All guides · Agent observability and evaluation

How to Keep Trace Data When an AI Agent Crashes Mid-Run

Part 9 of Agent observability and evaluation · 1,445 words · ~7 min read · published 2026-10-05 · updated 2026-10-05 · Markdown variant

How to keep trace data when an AI agent crashes mid-run: write a small run-start record outside the trace buffer, detect orphans, count them as unknown outcomes, and force-decide on shutdown.

In short

  • A buffered tail-sampling pipeline loses exactly the runs that crash, run out of memory or are killed, because the spans for those runs exist only in the process that died.
  • A run-start record written to storage outside the trace buffer, holding run ID, trace ID, versions, start time and parent ID, guarantees that every run leaves at least one durable trace of having begun.
  • An orphan, meaning a run with a start record and no close record, is a crash signal that no exception span can provide, since a dead process cannot report its own death.
  • Orphaned runs should be counted as an unscored unknown outcome rather than dropped, because excluding them makes online success rates look better than reality through survivorship bias.
  • As of October 2026 the reference corpus does not specify run-start records, heartbeats, orphan detection or flush-on-shutdown for agents, so this design is reasoned from the cited resources and no timing values are supplied.

Part of the AI Agent Observability and the Production Evaluation Playbook guide.


To keep trace data when an AI agent crashes mid-run, write a small run-start record to durable storage when the run begins, outside the in-memory trace buffer, and treat any start with no matching close as a crash signal. As of October 2026 the reference corpus does not specify run-start records, heartbeats, orphan detection or flush-on-shutdown, so what follows is this article's own reasoned design, not a documented standard. It extends one sentence in the sampling article and sits under the pillar. No interval, timeout or grace window is given here.

Why does a crash lose exactly the traces you want?

A crash loses the traces you want because a tail-based pipeline keeps each run's spans in memory until the run ends, and a dead process never reaches that point. The agent-observability resource describes a trace as one run under a stable ID, with errors and retries recorded on the span that failed. That model assumes the span gets exported. When the process is killed by an out-of-memory event, a deploy, a hard timeout or a host failure, the spans for the in-flight run are gone, and so is the exception that would have explained it.

The effect is selection in the wrong direction:

This is not a claim about any named tracing tool. Whether your exporter writes partial spans, flushes on exit or drops them is tool-specific, so read your own stack's documentation and test it by killing a run on purpose.

What goes into a run-start record?

A run-start record should hold only identity and context, never payloads: run ID, trace ID, version tags, start time and parent ID. It is written once, at the first moment the run exists, so it must be cheap and must not depend on the buffer.

Field Why it is there
Run ID The key the later close record and any orphan check match on
Trace ID Links the record to the full trace if one is later exported; the corpus describes one stable ID per run
Version tags Prompt, model and code versions, so a crash can be attributed; tagging itself is covered in rollout and rollback
Start time The only clock reading you are sure to have for a run that never finished
Parent ID The delegating run, so a crashed sub-agent can be tied to its tree

Keep the record small on purpose. No prompt text, tool arguments or user content belongs in it, which also keeps it clear of the redact-before-logging rule in the agent-observability resource. If a field would need redaction, it is too heavy for this record.

Where should the run-start record be written?

Write the run-start record to a store that outlives the agent process and is separate from the trace exporter, such as an append-only table, a queue or a key-value store. The point of separation is failure independence: if the exporter and the buffer die together, the record must not die with them.

Two properties matter more than the choice of store:

  1. Write before work. The record lands before the first model or tool call, so even an immediate crash leaves it.
  2. A second, small write at the end. A close record carries the outcome and duration. It can be a plain status update on the same row.

A mid-flight step may re-run after a restart, as the durable-execution-for-agents resource notes about replay. A start record therefore does not mean the step ran once, and it says nothing about side effects. If your runs resume on a durable engine, the engine's own history is the authority on what happened, and the start record is only an index of what began. Mechanics are in durable execution for agents.

How do you detect an orphaned run?

An orphaned run is a start record with no close record, and you detect it by periodically querying for starts that have stayed open longer than you are willing to wait. This is the one crash signal that works when the process cannot speak. An exception span needs a live process to be written, while an orphan is inferred from absence.

Two ways to age a start record into an orphan are available:

Heartbeats catch a crash earlier but add writes and a second thing that can fail; a deadline needs no extra writes but is slow to notice. Which you choose, and the values, are yours. No number is supplied here and the corpus supplies none.

An orphan is a signal and not a verdict. A run can be slow, paused for human approval, or still working, so route orphans to a person or an incident runbook rather than auto-classifying them as failures.

How should orphans appear in your metrics?

Orphans should appear as their own unknown outcome in online metrics, counted in the denominator and never scored as pass or fail. Survivorship bias is the reason: a success rate computed only over closed runs excludes every run that died, so crashes make the number go up.

A workable report shows, for each window:

The start-record count is the honest denominator. Denominator choice for online rates is covered in online agent metrics and drift monitoring; this article's contribution is the supply of a denominator that cannot lose runs. Show the unknown share beside every rate, and watch it per version tag, since a rise in orphans after a release is itself a regression signal.

Do not feed orphans to a judge or an evaluation set as if they had outputs. They have none. If a crash is later explained and reproduced, it can become a regression check through the incident-to-fixture loop.

What should happen when the process is told to stop?

When the process receives a shutdown signal, it should force-decide every in-flight run, flush what it has, and write close records marked as interrupted. A graceful stop is the one crash-like event you can act on, and the sampling article already recommends force-deciding a run that cannot finish.

A reasonable shutdown routine:

  1. Stop accepting further runs.
  2. For each in-flight run, apply the keep rule to the spans gathered by that point, exporting them as a partial trace if kept.
  3. Write a close record with outcome set to interrupted, distinct from error and from unknown.
  4. Exit.

Signals you cannot catch, such as a forced kill or a power loss, skip all of this, which is why the start record and orphan check exist as the backstop. The grace period a platform gives between a stop request and a forced kill is platform-specific, and no value is claimed here.

What does ChangeGamer do that resembles a start record?

ChangeGamer's article cycle pushes an empty commit named after the article before doing any work, which is an analogy and not an agent trace. Section 3a of .claude/commands/article-cycle.md says to create the branch and run git commit --allow-empty -m "Article cycle: start <slug>", then push it, immediately after the plan is returned and before writing.

The file's stated reason is that a pushed branch turns a session that dies mid-run into a resumable one: the next run finds the branch and continues. The commit holds no work, only the fact that work began, which is what a run-start record does. A branch is not a trace store, and ChangeGamer runs no agent trace store, so nothing here measures an agent system.

What the corpus does not tell you

The corpus does not specify the shape, storage, timing or orphan rules in this article, as of October 2026. The agent-observability resource supports the trace model and the redaction rule, and durable-execution-for-agents supports the replay caveat. The rest is reasoned design and should be treated as a hypothesis. Test it by killing a run on purpose and checking that the start record exists, that the orphan appears and that your success rate shows an unknown share. No number was supplied for any interval, timeout or grace window.

Frequently asked questions

Why do I lose traces when my AI agent crashes?
Traces are lost on a crash because tail-based pipelines hold a run in memory until it ends, so a process killed mid-run takes its buffered spans with it. The runs that crash are therefore the ones with no trace, which is the opposite of what debugging needs.
What is a run-start record for an AI agent?
A run-start record is a small durable entry written when an agent run begins, outside the trace buffer, holding fields such as run ID, trace ID, version tags, start time and parent ID. It proves the run existed even if the process dies before exporting any span.
How do I detect that an agent run crashed?
Detect a crashed agent run by finding start records that never received a matching close record. That orphan is the crash signal, because a process that was killed cannot write an error span. How long you wait before declaring an orphan is your own choice, since no number is published.
Should crashed runs count as failures in my success rate?
Crashed runs should be reported as a separate unknown outcome rather than silently dropped or forced into pass or fail. Leaving them out inflates the success rate, and calling them failures asserts something nobody scored. Show the unknown share beside every rate.

#agents #observability #tracing #reliability #crashes #metrics

Put this corpus inside your own agents

This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.

Agents: this page as Markdown · JSON · offers at /api/pricing.json · payment methods at /api/payment.json · single-resource access via HTTP 402 (how that works)