{
  "slug": "crash-safe-run-start-records-for-agent-traces",
  "title": "How to Keep Trace Data When an AI Agent Crashes Mid-Run",
  "description": "How to keep trace data when an AI agent crashes mid-run: write a small run-start record outside the trace buffer, detect orphans, count them as unknown outcomes, and force-decide on shutdown.",
  "kind": "sub",
  "order": 9,
  "target_query": "how to keep trace data when an AI agent crashes mid-run",
  "secondary_queries": [
    "ai agent observability",
    "agent run start record",
    "lost traces when an llm agent crashes"
  ],
  "tags": [
    "agents",
    "observability",
    "tracing",
    "reliability",
    "crashes",
    "metrics"
  ],
  "published": "2026-10-05",
  "updated": "2026-10-05",
  "words": 1445,
  "estimated_tokens": 1922,
  "premium": false,
  "rights": {
    "access": "free",
    "note": "Editorial guides are always free and never part of the licensed corpus.",
    "license": "https://changegamer.ai/license.xml",
    "pricing": "https://changegamer.ai/api/pricing.json",
    "payment": "https://changegamer.ai/api/payment.json"
  },
  "license": "https://changegamer.ai/license.xml",
  "citation": "ChangeGamer (2026-10-05). How to Keep Trace Data When an AI Agent Crashes Mid-Run. ChangeGamer. https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces (updated 2026-10-05).",
  "bibtex": "@misc{changegamer_crash_safe_run_start_records_for_agent_traces, title = {How to Keep Trace Data When an AI Agent Crashes Mid-Run}, publisher = {ChangeGamer}, year = {2026}, url = {https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces}, note = {Updated 2026-10-05}}",
  "canonical": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces",
  "markdown": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces.md",
  "takeaways": [
    "A buffered tail-sampling pipeline loses exactly the runs that crash, run out of memory or are killed, because the spans for those runs exist only in the process that died.",
    "A run-start record written to storage outside the trace buffer, holding run ID, trace ID, versions, start time and parent ID, guarantees that every run leaves at least one durable trace of having begun.",
    "An orphan, meaning a run with a start record and no close record, is a crash signal that no exception span can provide, since a dead process cannot report its own death.",
    "Orphaned runs should be counted as an unscored unknown outcome rather than dropped, because excluding them makes online success rates look better than reality through survivorship bias.",
    "As of October 2026 the reference corpus does not specify run-start records, heartbeats, orphan detection or flush-on-shutdown for agents, so this design is reasoned from the cited resources and no timing values are supplied."
  ],
  "outline": [
    {
      "depth": 2,
      "text": "Why does a crash lose exactly the traces you want?",
      "anchor": "why-does-a-crash-lose-exactly-the-traces-you-want",
      "url": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces#why-does-a-crash-lose-exactly-the-traces-you-want"
    },
    {
      "depth": 2,
      "text": "What goes into a run-start record?",
      "anchor": "what-goes-into-a-run-start-record",
      "url": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces#what-goes-into-a-run-start-record"
    },
    {
      "depth": 2,
      "text": "Where should the run-start record be written?",
      "anchor": "where-should-the-run-start-record-be-written",
      "url": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces#where-should-the-run-start-record-be-written"
    },
    {
      "depth": 2,
      "text": "How do you detect an orphaned run?",
      "anchor": "how-do-you-detect-an-orphaned-run",
      "url": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces#how-do-you-detect-an-orphaned-run"
    },
    {
      "depth": 2,
      "text": "How should orphans appear in your metrics?",
      "anchor": "how-should-orphans-appear-in-your-metrics",
      "url": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces#how-should-orphans-appear-in-your-metrics"
    },
    {
      "depth": 2,
      "text": "What should happen when the process is told to stop?",
      "anchor": "what-should-happen-when-the-process-is-told-to-stop",
      "url": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces#what-should-happen-when-the-process-is-told-to-stop"
    },
    {
      "depth": 2,
      "text": "What does ChangeGamer do that resembles a start record?",
      "anchor": "what-does-changegamer-do-that-resembles-a-start-record",
      "url": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces#what-does-changegamer-do-that-resembles-a-start-record"
    },
    {
      "depth": 2,
      "text": "What the corpus does not tell you",
      "anchor": "what-the-corpus-does-not-tell-you",
      "url": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces#what-the-corpus-does-not-tell-you"
    }
  ],
  "faq": [
    {
      "question": "Why do I lose traces when my AI agent crashes?",
      "answer": "Traces are lost on a crash because tail-based pipelines hold a run in memory until it ends, so a process killed mid-run takes its buffered spans with it. The runs that crash are therefore the ones with no trace, which is the opposite of what debugging needs."
    },
    {
      "question": "What is a run-start record for an AI agent?",
      "answer": "A run-start record is a small durable entry written when an agent run begins, outside the trace buffer, holding fields such as run ID, trace ID, version tags, start time and parent ID. It proves the run existed even if the process dies before exporting any span."
    },
    {
      "question": "How do I detect that an agent run crashed?",
      "answer": "Detect a crashed agent run by finding start records that never received a matching close record. That orphan is the crash signal, because a process that was killed cannot write an error span. How long you wait before declaring an orphan is your own choice, since no number is published."
    },
    {
      "question": "Should crashed runs count as failures in my success rate?",
      "answer": "Crashed runs should be reported as a separate unknown outcome rather than silently dropped or forced into pass or fail. Leaving them out inflates the success rate, and calling them failures asserts something nobody scored. Show the unknown share beside every rate."
    }
  ],
  "body": "To keep trace data when an AI agent crashes mid-run, write a small run-start record to durable storage when the run begins, outside the in-memory trace buffer, and treat any start with no matching close as a crash signal. As of October 2026 the reference corpus does not specify run-start records, heartbeats, orphan detection or flush-on-shutdown, so what follows is this article's own reasoned design, not a documented standard. It extends one sentence in [the sampling article](/articles/sampling-and-retention-of-agent-traces) and sits under [the pillar](/articles/agent-observability-and-evaluation). No interval, timeout or grace window is given here.\n\n## Why does a crash lose exactly the traces you want?\n\nA crash loses the traces you want because a tail-based pipeline keeps each run's spans in memory until the run ends, and a dead process never reaches that point. The agent-observability resource describes a trace as one run under a stable ID, with errors and retries recorded on the span that failed. That model assumes the span gets exported. When the process is killed by an out-of-memory event, a deploy, a hard timeout or a host failure, the spans for the in-flight run are gone, and so is the exception that would have explained it.\n\nThe effect is selection in the wrong direction:\n\n- Runs that finish are exported and appear in dashboards.\n- Runs that crash leave nothing, so they appear in no count.\n- The more often the agent crashes, the healthier its stored traces look.\n\nThis is not a claim about any named tracing tool. Whether your exporter writes partial spans, flushes on exit or drops them is tool-specific, so read your own stack's documentation and test it by killing a run on purpose.\n\n## What goes into a run-start record?\n\nA run-start record should hold only identity and context, never payloads: run ID, trace ID, version tags, start time and parent ID. It is written once, at the first moment the run exists, so it must be cheap and must not depend on the buffer.\n\n| Field | Why it is there |\n|---|---|\n| Run ID | The key the later close record and any orphan check match on |\n| Trace ID | Links the record to the full trace if one is later exported; the corpus describes one stable ID per run |\n| Version tags | Prompt, model and code versions, so a crash can be attributed; tagging itself is covered in [rollout and rollback](/articles/agent-rollout-and-rollback) |\n| Start time | The only clock reading you are sure to have for a run that never finished |\n| Parent ID | The delegating run, so a crashed sub-agent can be tied to its tree |\n\nKeep the record small on purpose. No prompt text, tool arguments or user content belongs in it, which also keeps it clear of the redact-before-logging rule in the agent-observability resource. If a field would need redaction, it is too heavy for this record.\n\n## Where should the run-start record be written?\n\nWrite the run-start record to a store that outlives the agent process and is separate from the trace exporter, such as an append-only table, a queue or a key-value store. The point of separation is failure independence: if the exporter and the buffer die together, the record must not die with them.\n\nTwo properties matter more than the choice of store:\n\n1. **Write before work.** The record lands before the first model or tool call, so even an immediate crash leaves it.\n2. **A second, small write at the end.** A close record carries the outcome and duration. It can be a plain status update on the same row.\n\nA mid-flight step may re-run after a restart, as the durable-execution-for-agents resource notes about replay. A start record therefore does not mean the step ran once, and it says nothing about side effects. If your runs resume on a durable engine, the engine's own history is the authority on what happened, and the start record is only an index of what began. Mechanics are in [durable execution for agents](/articles/durable-execution-for-ai-agents).\n\n## How do you detect an orphaned run?\n\nAn orphaned run is a start record with no close record, and you detect it by periodically querying for starts that have stayed open longer than you are willing to wait. This is the one crash signal that works when the process cannot speak. An exception span needs a live process to be written, while an orphan is inferred from absence.\n\nTwo ways to age a start record into an orphan are available:\n\n- **Deadline-based:** declare an orphan once the run has been open longer than its maximum allowed duration.\n- **Heartbeat-based:** have the run update a last-seen field periodically, and declare an orphan when the field goes stale.\n\nHeartbeats catch a crash earlier but add writes and a second thing that can fail; a deadline needs no extra writes but is slow to notice. Which you choose, and the values, are yours. No number is supplied here and the corpus supplies none.\n\nAn orphan is a signal and not a verdict. A run can be slow, paused for human approval, or still working, so route orphans to a person or an [incident runbook](/articles/agent-incident-response-runbooks) rather than auto-classifying them as failures.\n\n## How should orphans appear in your metrics?\n\nOrphans should appear as their own unknown outcome in online metrics, counted in the denominator and never scored as pass or fail. Survivorship bias is the reason: a success rate computed only over closed runs excludes every run that died, so crashes make the number go up.\n\nA workable report shows, for each window:\n\n- runs started (from start records, which crashes cannot erase)\n- runs closed with a scored outcome\n- runs closed with an error\n- orphans, labelled unknown\n\nThe start-record count is the honest denominator. Denominator choice for online rates is covered in [online agent metrics and drift monitoring](/articles/online-agent-metrics-and-drift-monitoring); this article's contribution is the supply of a denominator that cannot lose runs. Show the unknown share beside every rate, and watch it per version tag, since a rise in orphans after a release is itself a regression signal.\n\nDo not feed orphans to a judge or an evaluation set as if they had outputs. They have none. If a crash is later explained and reproduced, it can become a regression check through [the incident-to-fixture loop](/articles/incident-to-eval-fixture-loop).\n\n## What should happen when the process is told to stop?\n\nWhen the process receives a shutdown signal, it should force-decide every in-flight run, flush what it has, and write close records marked as interrupted. A graceful stop is the one crash-like event you can act on, and the sampling article already recommends force-deciding a run that cannot finish.\n\nA reasonable shutdown routine:\n\n1. Stop accepting further runs.\n2. For each in-flight run, apply the keep rule to the spans gathered by that point, exporting them as a partial trace if kept.\n3. Write a close record with outcome set to interrupted, distinct from error and from unknown.\n4. Exit.\n\nSignals you cannot catch, such as a forced kill or a power loss, skip all of this, which is why the start record and orphan check exist as the backstop. The grace period a platform gives between a stop request and a forced kill is platform-specific, and no value is claimed here.\n\n## What does ChangeGamer do that resembles a start record?\n\nChangeGamer's article cycle pushes an empty commit named after the article before doing any work, which is an analogy and not an agent trace. Section 3a of `.claude/commands/article-cycle.md` says to create the branch and run `git commit --allow-empty -m \"Article cycle: start <slug>\"`, then push it, immediately after the plan is returned and before writing.\n\nThe file's stated reason is that a pushed branch turns a session that dies mid-run into a resumable one: the next run finds the branch and continues. The commit holds no work, only the fact that work began, which is what a run-start record does. A branch is not a trace store, and ChangeGamer runs no agent trace store, so nothing here measures an agent system.\n\n## What the corpus does not tell you\n\nThe corpus does not specify the shape, storage, timing or orphan rules in this article, as of October 2026. The agent-observability resource supports the trace model and the redaction rule, and durable-execution-for-agents supports the replay caveat. The rest is reasoned design and should be treated as a hypothesis. Test it by killing a run on purpose and checking that the start record exists, that the orphan appears and that your success rate shows an unknown share. No number was supplied for any interval, timeout or grace window.",
  "cluster": {
    "id": "agent-observability-evaluation",
    "title": "Agent observability and evaluation",
    "description": "How to observe and evaluate an AI agent already live in production — tracing spans for tool calls and retrieval steps, judge-based screening versus human acceptance of live output, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures — not the pre-ship, CI-gating evaluation methodology already owned by evaluating-ai-agents-in-ci, and not the trace/span field mechanics already owned by agent-observability-for-reliability, both in the completed agent-reliability cluster.",
    "status": "complete",
    "pillar": {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "kind": "pillar",
      "order": 0,
      "html": "https://changegamer.ai/articles/agent-observability-and-evaluation",
      "markdown": "https://changegamer.ai/articles/agent-observability-and-evaluation.md",
      "json": "https://changegamer.ai/api/articles/agent-observability-and-evaluation.json"
    },
    "articles": [
      {
        "slug": "multi-agent-trace-propagation",
        "title": "Distributed Tracing for Multi-Agent AI Systems",
        "description": "How trace ID propagation actually works across a multi-agent handoff, why handoff and delegation need different trace shapes, and why an orchestrator-level trace can hide a failed leg of a fan-out.",
        "kind": "sub",
        "order": 1,
        "html": "https://changegamer.ai/articles/multi-agent-trace-propagation",
        "markdown": "https://changegamer.ai/articles/multi-agent-trace-propagation.md",
        "json": "https://changegamer.ai/api/articles/multi-agent-trace-propagation.json"
      },
      {
        "slug": "llm-as-judge-screening-in-production",
        "title": "Designing an LLM-as-Judge Pipeline for Production AI Agents",
        "description": "An operator playbook for screening live AI agent output with an LLM judge: a confidence/stakes routing architecture to a human queue, continuous live-traffic rubric design, and per-bias mitigations for position, verbosity, and self-preference.",
        "kind": "sub",
        "order": 2,
        "html": "https://changegamer.ai/articles/llm-as-judge-screening-in-production",
        "markdown": "https://changegamer.ai/articles/llm-as-judge-screening-in-production.md",
        "json": "https://changegamer.ai/api/articles/llm-as-judge-screening-in-production.json"
      },
      {
        "slug": "agent-cost-telemetry-in-production",
        "title": "How to Track AI Agent Costs in Production",
        "description": "How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.",
        "kind": "sub",
        "order": 3,
        "html": "https://changegamer.ai/articles/agent-cost-telemetry-in-production",
        "markdown": "https://changegamer.ai/articles/agent-cost-telemetry-in-production.md",
        "json": "https://changegamer.ai/api/articles/agent-cost-telemetry-in-production.json"
      },
      {
        "slug": "incident-to-eval-fixture-loop",
        "title": "How to Turn an AI Agent Incident into an Evaluation Test Case",
        "description": "How to turn an AI agent incident into an evaluation test case: freeze the trace, redact it, minimize it to the failing decision, label the expected outcome with a trusted oracle, prove it fails then passes, and retire it later.",
        "kind": "sub",
        "order": 4,
        "html": "https://changegamer.ai/articles/incident-to-eval-fixture-loop",
        "markdown": "https://changegamer.ai/articles/incident-to-eval-fixture-loop.md",
        "json": "https://changegamer.ai/api/articles/incident-to-eval-fixture-loop.json"
      },
      {
        "slug": "retrieval-attribution-logging-in-production",
        "title": "How to Log RAG Retrieval in Production for Debugging Agent Answers",
        "description": "How to log RAG retrieval in production so a wrong agent answer can be debugged: per-retrieval fields, what reached the context window versus what the model cited, privacy by hashes and IDs, and retention and sampling choices.",
        "kind": "sub",
        "order": 5,
        "html": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production",
        "markdown": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production.md",
        "json": "https://changegamer.ai/api/articles/retrieval-attribution-logging-in-production.json"
      },
      {
        "slug": "online-feedback-signals-for-ai-agents",
        "title": "How to Measure AI Agent Quality from Live User Feedback",
        "description": "How to measure AI agent quality from live user feedback: why explicit ratings are sparse and biased, how re-asks, abandonment and escalations mislead, and how to join each signal to a trace and route it to review.",
        "kind": "sub",
        "order": 6,
        "html": "https://changegamer.ai/articles/online-feedback-signals-for-ai-agents",
        "markdown": "https://changegamer.ai/articles/online-feedback-signals-for-ai-agents.md",
        "json": "https://changegamer.ai/api/articles/online-feedback-signals-for-ai-agents.json"
      },
      {
        "slug": "online-agent-metrics-and-drift-monitoring",
        "title": "How to Detect Quality Drift in a Production AI Agent",
        "description": "How to detect quality drift in a production AI agent: baseline aggregate signals, alert on a diff against the baseline, and separate a prompt, model or tool-version change from a shift in traffic mix.",
        "kind": "sub",
        "order": 7,
        "html": "https://changegamer.ai/articles/online-agent-metrics-and-drift-monitoring",
        "markdown": "https://changegamer.ai/articles/online-agent-metrics-and-drift-monitoring.md",
        "json": "https://changegamer.ai/api/articles/online-agent-metrics-and-drift-monitoring.json"
      },
      {
        "slug": "sampling-and-retention-of-agent-traces",
        "title": "How to Sample and Retain Production AI Agent Traces",
        "description": "How to sample and retain production AI agent traces: head vs tail sampling, keep-all-errors plus a random baseline, whole-trace decisions for multi-agent runs, retention tiers and redaction before the clock starts.",
        "kind": "sub",
        "order": 8,
        "html": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces",
        "markdown": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/sampling-and-retention-of-agent-traces.json"
      },
      {
        "slug": "crash-safe-run-start-records-for-agent-traces",
        "title": "How to Keep Trace Data When an AI Agent Crashes Mid-Run",
        "description": "How to keep trace data when an AI agent crashes mid-run: write a small run-start record outside the trace buffer, detect orphans, count them as unknown outcomes, and force-decide on shutdown.",
        "kind": "sub",
        "order": 9,
        "html": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces",
        "markdown": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/crash-safe-run-start-records-for-agent-traces.json"
      },
      {
        "slug": "choosing-an-llm-observability-backend",
        "title": "How to Choose an LLM Observability Platform for AI Agents",
        "description": "How to choose an LLM observability platform: decide on OTel-native ingestion, self-host versus cloud, export portability, redaction hooks and retention support before comparing vendors.",
        "kind": "sub",
        "order": 10,
        "html": "https://changegamer.ai/articles/choosing-an-llm-observability-backend",
        "markdown": "https://changegamer.ai/articles/choosing-an-llm-observability-backend.md",
        "json": "https://changegamer.ai/api/articles/choosing-an-llm-observability-backend.json"
      },
      {
        "slug": "redacting-sensitive-data-from-agent-traces",
        "title": "How to Redact PII From AI Agent Traces: Placement, Testing and Cleanup",
        "description": "How to redact PII from AI agent traces in practice: where the redactor sits in the pipeline, what to do per span field, how to test it with seeded fake PII, and how to clean up after a leak.",
        "kind": "sub",
        "order": 11,
        "html": "https://changegamer.ai/articles/redacting-sensitive-data-from-agent-traces",
        "markdown": "https://changegamer.ai/articles/redacting-sensitive-data-from-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/redacting-sensitive-data-from-agent-traces.json"
      }
    ]
  },
  "navigation": {
    "pillar": {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "kind": "pillar",
      "order": 0,
      "html": "https://changegamer.ai/articles/agent-observability-and-evaluation",
      "markdown": "https://changegamer.ai/articles/agent-observability-and-evaluation.md",
      "json": "https://changegamer.ai/api/articles/agent-observability-and-evaluation.json"
    },
    "previous": {
      "slug": "sampling-and-retention-of-agent-traces",
      "title": "How to Sample and Retain Production AI Agent Traces",
      "description": "How to sample and retain production AI agent traces: head vs tail sampling, keep-all-errors plus a random baseline, whole-trace decisions for multi-agent runs, retention tiers and redaction before the clock starts.",
      "kind": "sub",
      "order": 8,
      "html": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces",
      "markdown": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces.md",
      "json": "https://changegamer.ai/api/articles/sampling-and-retention-of-agent-traces.json"
    },
    "next": {
      "slug": "choosing-an-llm-observability-backend",
      "title": "How to Choose an LLM Observability Platform for AI Agents",
      "description": "How to choose an LLM observability platform: decide on OTel-native ingestion, self-host versus cloud, export portability, redaction hooks and retention support before comparing vendors.",
      "kind": "sub",
      "order": 10,
      "html": "https://changegamer.ai/articles/choosing-an-llm-observability-backend",
      "markdown": "https://changegamer.ai/articles/choosing-an-llm-observability-backend.md",
      "json": "https://changegamer.ai/api/articles/choosing-an-llm-observability-backend.json"
    }
  },
  "resources": [
    {
      "slug": "agent-observability",
      "html": "https://changegamer.ai/resources/agent-observability",
      "markdown": "https://changegamer.ai/resources/agent-observability.md",
      "json": "https://changegamer.ai/api/resources/agent-observability.json"
    },
    {
      "slug": "durable-execution-for-agents",
      "html": "https://changegamer.ai/resources/durable-execution-for-agents",
      "markdown": "https://changegamer.ai/resources/durable-execution-for-agents.md",
      "json": "https://changegamer.ai/api/resources/durable-execution-for-agents.json"
    }
  ]
}