{
  "slug": "multi-agent-trace-propagation",
  "title": "Distributed Tracing for Multi-Agent AI Systems",
  "description": "How trace ID propagation actually works across a multi-agent handoff, why handoff and delegation need different trace shapes, and why an orchestrator-level trace can hide a failed leg of a fan-out.",
  "kind": "sub",
  "order": 1,
  "target_query": "distributed tracing for multi-agent AI systems",
  "secondary_queries": [
    "opentelemetry genai semantic conventions",
    "llm tracing",
    "trace id propagation",
    "multi-agent observability"
  ],
  "tags": [
    "agents",
    "observability",
    "tracing",
    "multi-agent",
    "opentelemetry",
    "production"
  ],
  "published": "2026-09-27",
  "updated": "2026-09-27",
  "words": 1734,
  "estimated_tokens": 2306,
  "premium": false,
  "rights": {
    "access": "free",
    "note": "Editorial guides are always free and never part of the licensed corpus.",
    "license": "https://changegamer.ai/license.xml",
    "pricing": "https://changegamer.ai/api/pricing.json",
    "payment": "https://changegamer.ai/api/payment.json"
  },
  "license": "https://changegamer.ai/license.xml",
  "citation": "ChangeGamer (2026-09-27). Distributed Tracing for Multi-Agent AI Systems. ChangeGamer. https://changegamer.ai/articles/multi-agent-trace-propagation (updated 2026-09-27).",
  "bibtex": "@misc{changegamer_multi_agent_trace_propagation, title = {Distributed Tracing for Multi-Agent AI Systems}, publisher = {ChangeGamer}, year = {2026}, url = {https://changegamer.ai/articles/multi-agent-trace-propagation}, note = {Updated 2026-09-27}}",
  "canonical": "https://changegamer.ai/articles/multi-agent-trace-propagation",
  "markdown": "https://changegamer.ai/articles/multi-agent-trace-propagation.md",
  "takeaways": [
    "A multi-agent handoff that fails to forward the parent trace_id does not throw an error — it silently starts a second, disconnected trace, which is why propagation has to be verified directly rather than assumed from the absence of errors.",
    "A handoff transfers full control to the receiving agent while the calling agent stops, while delegation keeps the orchestrator in control and aggregates results — a distinction multi-agent-orchestration-patterns names as a concrete cross-cutting decision every multi-agent builder has to make.",
    "An orchestrator-level trace can look successful even when one leg of a fan-out failed underneath it, because a parent span only reports what the orchestrator itself chose to record rather than an independently audited status from every child.",
    "A hierarchical manager-of-managers topology multiplies the tracing challenge across as many tiers as the hierarchy is deep, since both coordination overhead and error-propagation difficulty grow with every additional layer of sub-orchestrators.",
    "A parallelization fan-out pattern needs its trace to record which branch's output the aggregator actually selected, not merely that every branch completed, or a wrong final answer becomes impossible to trace back to the branch that produced it.",
    "Multi-agent orchestration patterns names propagating a shared trace_id across every subagent call as a required cross-cutting concern, stating plainly that without it, cross-agent debugging is impossible."
  ],
  "outline": [
    {
      "depth": 2,
      "text": "How does trace ID propagation work across a multi-agent handoff?",
      "anchor": "how-does-trace-id-propagation-work-across-a-multi-agent-handoff",
      "url": "https://changegamer.ai/articles/multi-agent-trace-propagation#how-does-trace-id-propagation-work-across-a-multi-agent-handoff"
    },
    {
      "depth": 2,
      "text": "Handoff vs. delegation: why the distinction changes what a trace needs to capture",
      "anchor": "handoff-vs-delegation-why-the-distinction-changes-what-a-trace-needs-to-capture",
      "url": "https://changegamer.ai/articles/multi-agent-trace-propagation#handoff-vs-delegation-why-the-distinction-changes-what-a-trace-needs-to-capture"
    },
    {
      "depth": 2,
      "text": "Why can an orchestrator-level trace miss a failed leg of a fan-out?",
      "anchor": "why-can-an-orchestrator-level-trace-miss-a-failed-leg-of-a-fan-out",
      "url": "https://changegamer.ai/articles/multi-agent-trace-propagation#why-can-an-orchestrator-level-trace-miss-a-failed-leg-of-a-fan-out"
    },
    {
      "depth": 2,
      "text": "Tracing three multi-agent topologies: orchestrator-workers, hierarchical, and fan-out",
      "anchor": "tracing-three-multi-agent-topologies-orchestrator-workers-hierarchical-and-fan-out",
      "url": "https://changegamer.ai/articles/multi-agent-trace-propagation#tracing-three-multi-agent-topologies-orchestrator-workers-hierarchical-and-fan-out"
    },
    {
      "depth": 2,
      "text": "What is the most common way a multi-agent trace silently breaks?",
      "anchor": "what-is-the-most-common-way-a-multi-agent-trace-silently-breaks",
      "url": "https://changegamer.ai/articles/multi-agent-trace-propagation#what-is-the-most-common-way-a-multi-agent-trace-silently-breaks"
    },
    {
      "depth": 2,
      "text": "A trace-propagation checklist for multi-agent pipelines",
      "anchor": "a-trace-propagation-checklist-for-multi-agent-pipelines",
      "url": "https://changegamer.ai/articles/multi-agent-trace-propagation#a-trace-propagation-checklist-for-multi-agent-pipelines"
    },
    {
      "depth": 2,
      "text": "Sources and further reading",
      "anchor": "sources-and-further-reading",
      "url": "https://changegamer.ai/articles/multi-agent-trace-propagation#sources-and-further-reading"
    }
  ],
  "faq": [
    {
      "question": "What happens if a trace ID isn't propagated across an agent handoff?",
      "answer": "A dropped trace ID at a handoff does not produce an error message — the receiving agent's tracing library simply mints a fresh trace_id for what it treats as a brand-new run, so every span downstream of that handoff lands in the observability backend as a second, complete-looking trace with no field connecting it back to the request that spawned it."
    },
    {
      "question": "What is the difference between a handoff and delegation in a multi-agent system?",
      "answer": "A handoff transfers full control of a task to another agent and the calling agent stops entirely, while delegation keeps the orchestrator in control and aggregates the delegated agent's results before continuing — a distinction multi-agent-orchestration-patterns names explicitly because it changes whether a trace needs to model one continuous execution path or a branching parent-child tree."
    },
    {
      "question": "Why doesn't an orchestrator-level trace always show a failed sub-agent?",
      "answer": "An orchestrator-level trace does not always show a failed sub-agent because the parent span only reports what the orchestrator itself recorded about its children, and a subagent that returns a plausible-looking text response instead of an explicit structured failure signal can let the orchestrator aggregate a partial result and report overall success while one branch never actually completed correctly."
    },
    {
      "question": "Do OpenTelemetry's GenAI semantic conventions define how to trace a multi-agent handoff?",
      "answer": "OpenTelemetry's GenAI semantic conventions define the general trace and span vocabulary agent tracing runs on, but as of September 2026 they do not specify the mechanics of propagating a trace ID across a handoff to a different agent, process, or vendor, which is why a team building a multi-agent pipeline needs to verify propagation directly rather than assume the wire format guarantees it."
    }
  ],
  "body": "Multi-agent tracing gets a single sentence in [the pillar](/articles/agent-observability-and-evaluation): a shared identifier travels with a run into any sub-agent it hands work off to, which is what makes it possible to reconstruct one pipeline's activity instead of reading each agent's output as a disconnected feed. This article is the deep-dive that sentence points to: what happens when a handoff does or does not carry that identifier forward, why a handoff and a delegation need different trace shapes, why a fan-out's parent-level trace can look clean while one branch actually failed, and how the picture changes across the topologies builders actually reach for. For the underlying trace/span model — the OpenTelemetry `gen_ai.*` attribute names, the five-field-per-span checklist — see [what should an AI agent's observability system capture](/articles/agent-observability-for-reliability); this article extends that model to the multi-agent case.\n\n## How does trace ID propagation work across a multi-agent handoff?\n\nTrace ID propagation across a multi-agent handoff works by carrying the exact same `trace_id` value the parent run already has into the call that invokes the next agent, so every span the receiving agent produces nests under that one shared identifier instead of starting a trace of its own. [Multi-agent orchestration patterns](/resources/multi-agent-orchestration-patterns) names this directly: \"multi-agent runs require a shared trace_id propagated across all subagent calls. Without it, cross-agent debugging is impossible.\"\n\nA dropped trace ID does not fail loudly. Nothing throws, no request returns an error — the receiving agent's tracing library simply has no parent identifier to attach to, so it mints a fresh `trace_id` and opens what looks like a brand-new, complete run. What was actually one pipeline now exists in the observability backend as two disconnected records, and no query against either one will ever surface the other.\n\nWhichever channel carries a handoff needs its own explicit code to forward the identifier, since none do it by default: an in-process call needs the trace context passed as an explicit argument rather than assumed to travel automatically; a queued message needs the trace ID as a field in the payload itself, since a queue has no execution context to carry it implicitly; and an HTTP request to a separately hosted agent needs it as a header or body field the receiving service explicitly reads — a vendor-hosted third-party agent may not accept an inbound trace ID at all.\n\nNo resource in this site's corpus specifies which transport a given framework picks by default; the wire-level attribute names are covered in full in [what should an AI agent's observability system capture](/articles/agent-observability-for-reliability).\n\n## Handoff vs. delegation: why the distinction changes what a trace needs to capture\n\nA handoff and a delegation need different trace shapes because they are different control-flow patterns: multi-agent-orchestration-patterns defines a handoff as a transfer of full control where \"the calling agent stops,\" while delegation \"keeps the orchestrator in control and aggregates results.\" A trace built for one shape misreads the other.\n\nA handoff's trace needs to model one continuous execution path: control passes from agent A to agent B, and the run continues as B's problem. The parent span for A can close once the handoff fires — there is no moment where both agents are simultaneously \"in flight\" for the same task.\n\nA delegation's trace needs the opposite shape: the orchestrator's span stays open for the full duration of every child it spawned, not just the dispatch moment, and each delegated child needs its own span with an explicit terminal status, since the orchestrator's own status depends on collecting all of them. A partial return — nine children finished, one still running — is a representable state under delegation; a handoff's linear model has no way to express it, since a handoff only ever has one active branch at a time.\n\nReading multi-agent-orchestration-patterns' nine named patterns through this handoff-versus-delegation lens is this article's own extension, not a split that resource itself draws: routing reads as a sequence of handoffs, since a classifier step hands a task to one specialist and steps out, while orchestrator-workers, parallelization, and hierarchical manager-of-managers are delegation-shaped by design, since each keeps a coordinating agent in control while children run. Tooling built only for a handoff's linear model is why teams building an orchestrator-workers system on it discover their traces cannot represent a still-running child at all.\n\n## Why can an orchestrator-level trace miss a failed leg of a fan-out?\n\nAn orchestrator-level trace can miss a failed leg of a fan-out because the orchestrator's top-level span only reports what the orchestrator itself chose to record about its children, and a failing subagent that returns a plausible-looking text response instead of an explicit failure signal gives the orchestrator no independent way to know that branch broke. Multi-agent-orchestration-patterns names this directly: \"in a multi-agent pipeline, a subagent failure can silently corrupt downstream results,\" and the fix it prescribes is structural — \"subagents must return structured success/failure signals, not just text\" — not something a trace viewer can retrofit after the fact.\n\nConcretely: a ten-worker fan-out where one worker times out, or silently returns an empty result, can still let the orchestrator assemble what looks like a complete answer from the other nine. The aggregate response is wrong or incomplete, but nothing in the parent span's status field records that — a person reading only the orchestrator's span sees a successful run, while the actual failure sits one level down, in a child span whose status was never checked.\n\nThis is why a per-child structured status, not just a per-child span, is the load-bearing artifact: a fan-out's trace is only as trustworthy as its narrowest, most-ignored child. The resource's own guidance here — retry, degrade, or surface the gap — depends entirely on that per-child signal existing in the first place; without it, surfacing the gap has nothing concrete to surface.\n\n## Tracing three multi-agent topologies: orchestrator-workers, hierarchical, and fan-out\n\nTracing needs differ across multi-agent topologies because each one changes how many children a parent span tracks, how deep a viewer must render nesting, or which child's output made it into the final answer. Three of the nine patterns multi-agent-orchestration-patterns documents carry distinct observability implications, covered here at the observability angle only:\n\n| Topology | What changes for the trace | Observability requirement |\n|---|---|---|\n| Orchestrator-workers | Worker count is decided at runtime, not fixed in advance | Trace schema must support a variable-cardinality set of children per parent span |\n| Hierarchical / manager-of-managers | The pattern repeats across tiers, so a bottom-tier failure surfaces through several intermediate parents | Viewer must render N levels of nesting; the resource flags that \"error propagation is harder to trace\" and \"observability becomes critical\" at this depth |\n| Parallelization (fan-out / voting) | Every branch runs the same task; only one branch's output — or a merge — reaches the caller | Trace needs a field recording which branch was selected or how the merge was computed, not just that every branch completed |\n\nOrchestrator-workers and hierarchical topologies share one pressure: as tiers multiply, so do the places a trace ID could fail to propagate, since every extra layer of sub-orchestrators is one more boundary it has to cross correctly. Parallelization has the opposite shape — propagation across branches is usually fine, since every branch shares one dispatch point — but without a selection field, a wrong answer traces to which branch ran, not to why the aggregator picked that branch's output.\n\n## What is the most common way a multi-agent trace silently breaks?\n\nThe most common way a multi-agent trace silently breaks, as of September 2026, is a handoff that forwards everything except the trace ID itself — the receiving agent gets the right task and context and still opens a disconnected trace, because propagation was never part of the interface contract the two agents agreed on. No resource in this site's corpus names this specific failure mode; treat it as a reasoned inference from how propagation mechanically works, not a documented industry finding.\n\nThree situations produce it repeatedly: a framework upgrade that changes how context is threaded internally and quietly stops forwarding a custom trace field; an async or queue boundary that loses whatever in-memory context carried the trace ID, since a queue has no built-in execution context to preserve; and a third-party or vendor-hosted agent that was never built to accept an inbound trace ID at all, so a caller that forwards one correctly still has nowhere for it to land.\n\nBecause none of these failures throw an error, the way to catch them is to watch the aggregate shape of the trace store, not any one trace: a rising count of root-level traces — no parent, starting fresh — that should have arrived as a child of a known entry point is the signal a broken propagation path produces.\n\n## A trace-propagation checklist for multi-agent pipelines\n\nBefore treating a multi-agent pipeline's tracing as complete, it should clear this list:\n\n- Every handoff and delegation call forwards the parent trace_id explicitly, verified across at least one process, queue, or vendor boundary — not just confirmed in-process, where propagation is easiest and least likely to break.\n- The trace model distinguishes a handoff's linear continuation from a delegation's branching aggregation, matching whichever pattern the pipeline actually uses.\n- Every subagent returns a structured success/failure signal on its own span, so an orchestrator's aggregate status can never silently mask one child's failure.\n- A hierarchical topology's trace viewer renders as many nesting tiers as the manager-of-managers hierarchy is actually deep.\n- A parallelization or fan-out pipeline records which branch's output was selected or how branches were merged, not only that every branch finished.\n- The trace store is monitored for an unexplained rise in root-level, parentless traces — that aggregate signal, not any single trace, is what a silently broken propagation path actually produces.\n\n## Sources and further reading\n\nEvery claim above traces back to one of two corpus resources:\n\n- Handoff-vs-delegation, partial subagent failure, structured success/failure signals, shared trace_id propagation: [/resources/multi-agent-orchestration-patterns](/resources/multi-agent-orchestration-patterns)\n- The base trace/span model this article extends: [/resources/agent-observability](/resources/agent-observability)\n\nFor the wire-level attribute names, see [what should an AI agent's observability system capture](/articles/agent-observability-for-reliability). For where this fits the full observability-and-evaluation surface, see [the pillar](/articles/agent-observability-and-evaluation). ChangeGamer does not run a multi-agent pipeline itself as of September 2026, so no first-hand operational example appears above — every mechanism here is grounded in a corpus resource or flagged as a reasoned inference.",
  "cluster": {
    "id": "agent-observability-evaluation",
    "title": "Agent observability and evaluation",
    "description": "How to observe and evaluate an AI agent already live in production — tracing spans for tool calls and retrieval steps, judge-based screening versus human acceptance of live output, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures — not the pre-ship, CI-gating evaluation methodology already owned by evaluating-ai-agents-in-ci, and not the trace/span field mechanics already owned by agent-observability-for-reliability, both in the completed agent-reliability cluster.",
    "status": "complete",
    "pillar": {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "kind": "pillar",
      "order": 0,
      "html": "https://changegamer.ai/articles/agent-observability-and-evaluation",
      "markdown": "https://changegamer.ai/articles/agent-observability-and-evaluation.md",
      "json": "https://changegamer.ai/api/articles/agent-observability-and-evaluation.json"
    },
    "articles": [
      {
        "slug": "multi-agent-trace-propagation",
        "title": "Distributed Tracing for Multi-Agent AI Systems",
        "description": "How trace ID propagation actually works across a multi-agent handoff, why handoff and delegation need different trace shapes, and why an orchestrator-level trace can hide a failed leg of a fan-out.",
        "kind": "sub",
        "order": 1,
        "html": "https://changegamer.ai/articles/multi-agent-trace-propagation",
        "markdown": "https://changegamer.ai/articles/multi-agent-trace-propagation.md",
        "json": "https://changegamer.ai/api/articles/multi-agent-trace-propagation.json"
      },
      {
        "slug": "llm-as-judge-screening-in-production",
        "title": "Designing an LLM-as-Judge Pipeline for Production AI Agents",
        "description": "An operator playbook for screening live AI agent output with an LLM judge: a confidence/stakes routing architecture to a human queue, continuous live-traffic rubric design, and per-bias mitigations for position, verbosity, and self-preference.",
        "kind": "sub",
        "order": 2,
        "html": "https://changegamer.ai/articles/llm-as-judge-screening-in-production",
        "markdown": "https://changegamer.ai/articles/llm-as-judge-screening-in-production.md",
        "json": "https://changegamer.ai/api/articles/llm-as-judge-screening-in-production.json"
      },
      {
        "slug": "agent-cost-telemetry-in-production",
        "title": "How to Track AI Agent Costs in Production",
        "description": "How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.",
        "kind": "sub",
        "order": 3,
        "html": "https://changegamer.ai/articles/agent-cost-telemetry-in-production",
        "markdown": "https://changegamer.ai/articles/agent-cost-telemetry-in-production.md",
        "json": "https://changegamer.ai/api/articles/agent-cost-telemetry-in-production.json"
      },
      {
        "slug": "incident-to-eval-fixture-loop",
        "title": "How to Turn an AI Agent Incident into an Evaluation Test Case",
        "description": "How to turn an AI agent incident into an evaluation test case: freeze the trace, redact it, minimize it to the failing decision, label the expected outcome with a trusted oracle, prove it fails then passes, and retire it later.",
        "kind": "sub",
        "order": 4,
        "html": "https://changegamer.ai/articles/incident-to-eval-fixture-loop",
        "markdown": "https://changegamer.ai/articles/incident-to-eval-fixture-loop.md",
        "json": "https://changegamer.ai/api/articles/incident-to-eval-fixture-loop.json"
      },
      {
        "slug": "retrieval-attribution-logging-in-production",
        "title": "How to Log RAG Retrieval in Production for Debugging Agent Answers",
        "description": "How to log RAG retrieval in production so a wrong agent answer can be debugged: per-retrieval fields, what reached the context window versus what the model cited, privacy by hashes and IDs, and retention and sampling choices.",
        "kind": "sub",
        "order": 5,
        "html": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production",
        "markdown": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production.md",
        "json": "https://changegamer.ai/api/articles/retrieval-attribution-logging-in-production.json"
      },
      {
        "slug": "online-feedback-signals-for-ai-agents",
        "title": "How to Measure AI Agent Quality from Live User Feedback",
        "description": "How to measure AI agent quality from live user feedback: why explicit ratings are sparse and biased, how re-asks, abandonment and escalations mislead, and how to join each signal to a trace and route it to review.",
        "kind": "sub",
        "order": 6,
        "html": "https://changegamer.ai/articles/online-feedback-signals-for-ai-agents",
        "markdown": "https://changegamer.ai/articles/online-feedback-signals-for-ai-agents.md",
        "json": "https://changegamer.ai/api/articles/online-feedback-signals-for-ai-agents.json"
      },
      {
        "slug": "online-agent-metrics-and-drift-monitoring",
        "title": "How to Detect Quality Drift in a Production AI Agent",
        "description": "How to detect quality drift in a production AI agent: baseline aggregate signals, alert on a diff against the baseline, and separate a prompt, model or tool-version change from a shift in traffic mix.",
        "kind": "sub",
        "order": 7,
        "html": "https://changegamer.ai/articles/online-agent-metrics-and-drift-monitoring",
        "markdown": "https://changegamer.ai/articles/online-agent-metrics-and-drift-monitoring.md",
        "json": "https://changegamer.ai/api/articles/online-agent-metrics-and-drift-monitoring.json"
      },
      {
        "slug": "sampling-and-retention-of-agent-traces",
        "title": "How to Sample and Retain Production AI Agent Traces",
        "description": "How to sample and retain production AI agent traces: head vs tail sampling, keep-all-errors plus a random baseline, whole-trace decisions for multi-agent runs, retention tiers and redaction before the clock starts.",
        "kind": "sub",
        "order": 8,
        "html": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces",
        "markdown": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/sampling-and-retention-of-agent-traces.json"
      },
      {
        "slug": "crash-safe-run-start-records-for-agent-traces",
        "title": "How to Keep Trace Data When an AI Agent Crashes Mid-Run",
        "description": "How to keep trace data when an AI agent crashes mid-run: write a small run-start record outside the trace buffer, detect orphans, count them as unknown outcomes, and force-decide on shutdown.",
        "kind": "sub",
        "order": 9,
        "html": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces",
        "markdown": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/crash-safe-run-start-records-for-agent-traces.json"
      },
      {
        "slug": "choosing-an-llm-observability-backend",
        "title": "How to Choose an LLM Observability Platform for AI Agents",
        "description": "How to choose an LLM observability platform: decide on OTel-native ingestion, self-host versus cloud, export portability, redaction hooks and retention support before comparing vendors.",
        "kind": "sub",
        "order": 10,
        "html": "https://changegamer.ai/articles/choosing-an-llm-observability-backend",
        "markdown": "https://changegamer.ai/articles/choosing-an-llm-observability-backend.md",
        "json": "https://changegamer.ai/api/articles/choosing-an-llm-observability-backend.json"
      },
      {
        "slug": "redacting-sensitive-data-from-agent-traces",
        "title": "How to Redact PII From AI Agent Traces: Placement, Testing and Cleanup",
        "description": "How to redact PII from AI agent traces in practice: where the redactor sits in the pipeline, what to do per span field, how to test it with seeded fake PII, and how to clean up after a leak.",
        "kind": "sub",
        "order": 11,
        "html": "https://changegamer.ai/articles/redacting-sensitive-data-from-agent-traces",
        "markdown": "https://changegamer.ai/articles/redacting-sensitive-data-from-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/redacting-sensitive-data-from-agent-traces.json"
      }
    ]
  },
  "navigation": {
    "pillar": {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "kind": "pillar",
      "order": 0,
      "html": "https://changegamer.ai/articles/agent-observability-and-evaluation",
      "markdown": "https://changegamer.ai/articles/agent-observability-and-evaluation.md",
      "json": "https://changegamer.ai/api/articles/agent-observability-and-evaluation.json"
    },
    "previous": null,
    "next": {
      "slug": "llm-as-judge-screening-in-production",
      "title": "Designing an LLM-as-Judge Pipeline for Production AI Agents",
      "description": "An operator playbook for screening live AI agent output with an LLM judge: a confidence/stakes routing architecture to a human queue, continuous live-traffic rubric design, and per-bias mitigations for position, verbosity, and self-preference.",
      "kind": "sub",
      "order": 2,
      "html": "https://changegamer.ai/articles/llm-as-judge-screening-in-production",
      "markdown": "https://changegamer.ai/articles/llm-as-judge-screening-in-production.md",
      "json": "https://changegamer.ai/api/articles/llm-as-judge-screening-in-production.json"
    }
  },
  "resources": [
    {
      "slug": "multi-agent-orchestration-patterns",
      "html": "https://changegamer.ai/resources/multi-agent-orchestration-patterns",
      "markdown": "https://changegamer.ai/resources/multi-agent-orchestration-patterns.md",
      "json": "https://changegamer.ai/api/resources/multi-agent-orchestration-patterns.json"
    },
    {
      "slug": "agent-observability",
      "html": "https://changegamer.ai/resources/agent-observability",
      "markdown": "https://changegamer.ai/resources/agent-observability.md",
      "json": "https://changegamer.ai/api/resources/agent-observability.json"
    }
  ]
}