{
  "slug": "llm-as-judge-screening-in-production",
  "title": "Designing an LLM-as-Judge Pipeline for Production AI Agents",
  "description": "An operator playbook for screening live AI agent output with an LLM judge: a confidence/stakes routing architecture to a human queue, continuous live-traffic rubric design, and per-bias mitigations for position, verbosity, and self-preference.",
  "kind": "sub",
  "order": 2,
  "target_query": "designing an LLM-as-judge pipeline for production AI agents",
  "secondary_queries": [
    "ai agent evaluation",
    "llm judge routing architecture",
    "judge confidence scoring",
    "human review queue for ai agent output",
    "rubric-based grading for ai agents",
    "llm as judge bias mitigation"
  ],
  "tags": [
    "agents",
    "evaluation",
    "llm-as-judge",
    "observability",
    "production",
    "rubric"
  ],
  "published": "2026-09-28",
  "updated": "2026-09-28",
  "words": 1620,
  "estimated_tokens": 2155,
  "premium": false,
  "rights": {
    "access": "free",
    "note": "Editorial guides are always free and never part of the licensed corpus.",
    "license": "https://changegamer.ai/license.xml",
    "pricing": "https://changegamer.ai/api/pricing.json",
    "payment": "https://changegamer.ai/api/payment.json"
  },
  "license": "https://changegamer.ai/license.xml",
  "citation": "ChangeGamer (2026-09-28). Designing an LLM-as-Judge Pipeline for Production AI Agents. ChangeGamer. https://changegamer.ai/articles/llm-as-judge-screening-in-production (updated 2026-09-28).",
  "bibtex": "@misc{changegamer_llm_as_judge_screening_in_production, title = {Designing an LLM-as-Judge Pipeline for Production AI Agents}, publisher = {ChangeGamer}, year = {2026}, url = {https://changegamer.ai/articles/llm-as-judge-screening-in-production}, note = {Updated 2026-09-28}}",
  "canonical": "https://changegamer.ai/articles/llm-as-judge-screening-in-production",
  "markdown": "https://changegamer.ai/articles/llm-as-judge-screening-in-production.md",
  "takeaways": [
    "A confidence-and-stakes routing architecture scores a judge's verdict on two separate axes — how confident the judge itself was and how costly a wrong answer would be — rather than collapsing both into one pass/fail threshold.",
    "Disagreement across repeated judge calls on the identical output is a more measurable low-confidence signal than a judge's own self-reported uncertainty, since a self-report is itself just another claim the same model is making.",
    "A human review queue for judge-flagged AI agent output works best as a priority-ranked list rather than a first-in-first-out inbox, so a flag combining low confidence with high stakes reaches a reviewer before a flag carrying only one of the two.",
    "Position bias, verbosity bias, and self-preference bias each have a named mitigation in the evaluating-ai-agents resource: randomize and average flipped pairwise order, reward concision in the rubric, and score with a judge from a different model family than the production model.",
    "ChangeGamer's own scripts/aeo-audit.mjs scores a page's structural quotability through fixed, deterministic heuristics with zero network or model calls, and it does not grade whether any claim on the page is actually factually correct.",
    "A live rubric that screens continuous production traffic needs its own version stamp, since comparing a run scored under a since-revised rubric to a run scored under the current one measures two different things as if they were one."
  ],
  "outline": [
    {
      "depth": 2,
      "text": "A confidence-and-stakes routing architecture for live judge screening",
      "anchor": "a-confidence-and-stakes-routing-architecture-for-live-judge-screening",
      "url": "https://changegamer.ai/articles/llm-as-judge-screening-in-production#a-confidence-and-stakes-routing-architecture-for-live-judge-screening"
    },
    {
      "depth": 2,
      "text": "What counts as a low-confidence judge verdict?",
      "anchor": "what-counts-as-a-low-confidence-judge-verdict",
      "url": "https://changegamer.ai/articles/llm-as-judge-screening-in-production#what-counts-as-a-low-confidence-judge-verdict"
    },
    {
      "depth": 2,
      "text": "What does a human review queue actually look like operationally?",
      "anchor": "what-does-a-human-review-queue-actually-look-like-operationally",
      "url": "https://changegamer.ai/articles/llm-as-judge-screening-in-production#what-does-a-human-review-queue-actually-look-like-operationally"
    },
    {
      "depth": 2,
      "text": "Designing a rubric for continuous live-traffic screening",
      "anchor": "designing-a-rubric-for-continuous-live-traffic-screening",
      "url": "https://changegamer.ai/articles/llm-as-judge-screening-in-production#designing-a-rubric-for-continuous-live-traffic-screening"
    },
    {
      "depth": 2,
      "text": "Mitigating position, verbosity, and self-preference bias in a live judge",
      "anchor": "mitigating-position-verbosity-and-self-preference-bias-in-a-live-judge",
      "url": "https://changegamer.ai/articles/llm-as-judge-screening-in-production#mitigating-position-verbosity-and-self-preference-bias-in-a-live-judge"
    },
    {
      "depth": 2,
      "text": "Does this architecture apply to pre-ship CI gating too?",
      "anchor": "does-this-architecture-apply-to-pre-ship-ci-gating-too",
      "url": "https://changegamer.ai/articles/llm-as-judge-screening-in-production#does-this-architecture-apply-to-pre-ship-ci-gating-too"
    },
    {
      "depth": 2,
      "text": "Sources and further reading",
      "anchor": "sources-and-further-reading",
      "url": "https://changegamer.ai/articles/llm-as-judge-screening-in-production#sources-and-further-reading"
    }
  ],
  "faq": [
    {
      "question": "How does a judge's low-confidence flag get prioritized against a high-stakes flag in a human review queue?",
      "answer": "One reasoned design multiplies a confidence-uncertainty score — how much repeated judge calls on the same output disagreed — by a stakes weight tied to the task, then ranks the review queue by that combined product, so a verdict that is both uncertain and costly to get wrong reaches a reviewer before a verdict carrying only one of the two properties."
    },
    {
      "question": "What makes a live LLM-as-judge verdict count as low confidence?",
      "answer": "A verdict counts as low confidence when repeated judge calls against the identical output disagree with each other, or when the judge itself reports low certainty alongside its score, with disagreement across repeated calls treated as the sturdier of the two signals since a self-reported confidence level is itself just another claim from the same model."
    },
    {
      "question": "How is a live-traffic rubric for judge screening different from a fixed offline eval rubric?",
      "answer": "A live-traffic rubric has to survive real user interactions an offline rubric, built once against a fixed dataset, was never written against, which is why evaluating-ai-agents credits online evaluation specifically with catching distribution shift and contamination that offline evaluation cannot, and why a continuously screened rubric needs its own version stamp to keep differently-scored runs from being compared as if they measured the same thing."
    },
    {
      "question": "Can the same model mitigate self-preference bias by judging its own output?",
      "answer": "No — evaluating-ai-agents names using a different model family as judge as the specific mitigation for self-preference bias, precisely because a model rates its own outputs more favorably, so a production pipeline that reuses its serving model as its judge keeps the exact bias that mitigation exists to remove."
    },
    {
      "question": "Does ChangeGamer's aeo-audit.mjs score AI agent output the same way a live LLM judge does?",
      "answer": "No — aeo-audit.mjs is a deterministic, rule-based scorer that reads only a page's already-built markdown variant and its HTML output, with zero network or model calls, and checks fixed structural heuristics like answer-first-sentence length and question-heading ratio, while a live LLM-as-judge makes a probabilistic call about a response's actual quality; the audit script scores quotability and structure, never whether a claim on the page is factually correct."
    }
  ],
  "body": "A judge model screening live AI agent traffic produces a verdict on every sampled run, but a verdict is not by itself a decision about what happens next — that takes an architecture, not a threshold. [The pillar](/articles/agent-observability-and-evaluation) sketches this in a single hedged aside, inferred from how narrowly evaluating-ai-agents scopes human review: send the judge's shakiest calls and its costliest-to-miss ones somewhere a person can catch them, and let the rest move on untouched. This article turns that aside into an actual pipeline — what makes a verdict count as shaky, what a human queue does with a flag once it lands, how a rubric should be written for continuous live traffic instead of a fixed offline set, and how to apply the corpus's own per-bias mitigations to that live case. Everything below is this article's own reasoned design extending that single aside, not a documented industry standard: no resource in this site's corpus specifies a production screening architecture at this depth, so treat the mechanism as one workable design, not the only one.\n\n## A confidence-and-stakes routing architecture for live judge screening\n\nA confidence-and-stakes routing architecture scores every judge verdict along two independent axes — how confident the judge itself was, and how costly a wrong answer would be — and only a verdict that lands low on both axes passes through unreviewed. Collapsing the two axes into one number is the mistake this architecture avoids: a judge can be highly confident and still wrong on an expensive run, and a low-confidence verdict on a genuinely low-stakes run does not justify a scarce reviewer's attention. The next two sections cover each axis: what produces the confidence signal, and how the combined score becomes queue order rather than a flat accept/reject gate.\n\n## What counts as a low-confidence judge verdict?\n\nA judge verdict counts as low-confidence when either of two measurable signals says so: the judge's own stated uncertainty, or disagreement across repeated calls on the identical output. A judge prompted to report a confidence level alongside its score produces a self-reported signal — worth capturing, but itself just another claim the same model is making, no more inherently reliable than the verdict it accompanies. A sturdier signal comes from calling the judge multiple times against the same output and checking whether the verdicts agree: an output that scores \"pass\" four times out of five carries real uncertainty a single call would have hidden completely. This design treats agreement across repeated calls as the primary confidence signal and self-reported uncertainty as a secondary, corroborating one — a choice this article is making, not something evaluating-ai-agents prescribes for judge screening specifically, though the resource's own distinction between a single pass/fail measurement and a multi-trial reliability metric is exactly the reasoning this choice extends from the agent being scored to the judge doing the scoring.\n\n## What does a human review queue actually look like operationally?\n\nA human review queue for judge-flagged output operates as a priority-ranked list, not a first-in-first-out inbox, because a flag combining low judge confidence with a high-stakes tag needs attention before a flag carrying only one of the two properties. One workable design: multiply a confidence-uncertainty score — near 0 when repeated judge calls fully agreed, near 1 when they fully split — by a stakes weight assigned per task type, and rank the queue by that product, so a run both unsure about and costly to miss surfaces first, regardless of arrival order.\n\nA run's own trace data can feed the stakes side of that score directly. Agent-observability documents errors and retries, with the original exception attached to the failing span, as signals a well-instrumented trace already captures; a run whose trace shows a tool-call error or a mid-execution retry is a reasonable candidate for a higher stakes weight than a clean run, on the theory that a recovery already happened once and the final output deserves a second look independent of what the judge scored it.\n\nTwo operational details this design still has to answer beyond scoring:\n\n- **Queue overflow** — a queue depth past what a reviewer pool can clear in a shift needs an escalation path, such as a second reviewer pool or a temporarily wider pass-through threshold, rather than letting flagged output age unreviewed.\n- **Feedback into the rubric** — every human decision that overturns a judge's verdict is a labeled example worth feeding back into the rubric the next section covers, closing a loop between what a person catches live and what the judge is scored against going forward.\n\n## Designing a rubric for continuous live-traffic screening\n\nA rubric for continuous live-traffic screening differs from a fixed offline eval rubric in what it has to survive, not in how it is written. Evaluating-ai-agents draws the underlying line directly: offline evaluation runs against \"a fixed dataset with pre-computed reference outputs,\" while online evaluation runs \"on live tasks drawn from real user interactions or a live environment\" specifically because live traffic \"catches distribution shift and contamination\" a fixed set structurally cannot. A rubric built once against a fixed set inherits that blind spot the moment it starts scoring live traffic it was never written against.\n\nThe same resource's general rubric guidance still holds online: \"define explicit scoring criteria before running the eval,\" and \"apply the rubric consistently... rubric-based grading reduces judge variance and produces auditable scores.\" What a continuously screened rubric additionally needs, in this design, is a version stamp — a start date, and an end date once a revision replaces it — so a run scored under one rubric version is never silently compared against a run scored under a different one as if the two measured the same thing.\n\nA live rubric that is never revised is not actually continuous, just an offline rubric applied to a moving stream. The trigger for a revision here is the same human-queue overturn data from the previous section: a cluster of overturns sharing a task type or a failure pattern the current rubric doesn't score for is the signal that criterion needs adding, not a fixed schedule detached from what live traffic is actually doing.\n\n## Mitigating position, verbosity, and self-preference bias in a live judge\n\nMitigating a live judge's bias means applying the same three fixes documented for offline LLM-as-judge scoring, adjusted for the cost constraint production adds. Evaluating-ai-agents names a specific mitigation for each of the three failure modes:\n\n- **Position bias** — \"randomizing order and averaging flipped-order results\" for any pairwise comparison the judge runs.\n- **Verbosity bias** — a rubric \"that reward[s] concision and penalize[s] padding,\" rather than one that implicitly rewards a longer, more formal-sounding answer.\n- **Self-preference** — \"using a different family as judge\" than the model family generating production output.\n\nAll three transfer to a live judge unmodified, but production adds a cost pressure offline scoring rarely faces: running a pairwise comparison twice, in both orders, doubles the judge calls, latency, and spend behind every comparison screened. The routing architecture above absorbs that cost: reserve flipped-order averaging for comparisons the stakes score already flagged as worth the second call, rather than doubling every judge call across all sampled traffic.\n\nSelf-preference deserves one production-specific note: a team already paying for one model family's API as its production model has an operational incentive to reach for that same family as its judge, and that convenience is exactly the setup the bias favors — a different family costs a second API relationship, but it is the only way the mitigation actually holds.\n\nWorth naming precisely what a live LLM-as-judge is not: it is a probabilistic scorer whose verdict can differ run to run on identical input, which is the entire reason the confidence signal above exists. Contrast that with a genuinely deterministic scorer: this site's own `scripts/aeo-audit.mjs`, as of September 2026, reads only a page's already-built markdown variant and HTML output, makes no network or model call, and scores fixed structural heuristics — how many words precede a section's answer, what share of headings are phrased as questions, whether a takeaway opens with a back-reference, how long an unbroken section runs, whether a time-sensitive claim carries a date. The same input always produces the same score, with no position, verbosity, or self-preference bias to mitigate in the first place. But that determinism buys quotability, not correctness: the audit cannot tell a well-structured false claim from a well-structured true one — exactly the judgment call a live judge, biases and all, exists to make instead.\n\n## Does this architecture apply to pre-ship CI gating too?\n\nNo — this architecture is scoped to live production traffic, not the pre-ship gate a release clears before it reaches users. [How to evaluate AI agents in CI](/articles/evaluating-ai-agents-in-ci) covers the identical bias set — position, verbosity, self-preference — for the case where a judge scores a candidate release against a fixed pre-ship suite rather than a continuous stream of live output; that is the article to read for that side of the boundary, and this one does not restate its treatment. The operational distinction: a CI gate blocks a release from shipping at all, so a false negative there costs a delayed deploy, while a live-screening miss ships straight to a real user before anyone catches it — the asymmetry the confidence-and-stakes routing architecture above exists to manage.\n\n## Sources and further reading\n\nThe LLM-as-judge biases, online-vs-offline distinction, and rubric-design guidance above are documented in [evaluating AI agents](/resources/evaluating-ai-agents). The trace-level error-and-retry signal feeding the stakes score is in [agent observability and tracing](/resources/agent-observability). For the pre-ship, CI-gating treatment of the same bias set, see [how to evaluate AI agents in CI](/articles/evaluating-ai-agents-in-ci). For where live judge screening fits the full production observability-and-evaluation surface, see [the pillar](/articles/agent-observability-and-evaluation).",
  "cluster": {
    "id": "agent-observability-evaluation",
    "title": "Agent observability and evaluation",
    "description": "How to observe and evaluate an AI agent already live in production — tracing spans for tool calls and retrieval steps, judge-based screening versus human acceptance of live output, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures — not the pre-ship, CI-gating evaluation methodology already owned by evaluating-ai-agents-in-ci, and not the trace/span field mechanics already owned by agent-observability-for-reliability, both in the completed agent-reliability cluster.",
    "status": "complete",
    "pillar": {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "kind": "pillar",
      "order": 0,
      "html": "https://changegamer.ai/articles/agent-observability-and-evaluation",
      "markdown": "https://changegamer.ai/articles/agent-observability-and-evaluation.md",
      "json": "https://changegamer.ai/api/articles/agent-observability-and-evaluation.json"
    },
    "articles": [
      {
        "slug": "multi-agent-trace-propagation",
        "title": "Distributed Tracing for Multi-Agent AI Systems",
        "description": "How trace ID propagation actually works across a multi-agent handoff, why handoff and delegation need different trace shapes, and why an orchestrator-level trace can hide a failed leg of a fan-out.",
        "kind": "sub",
        "order": 1,
        "html": "https://changegamer.ai/articles/multi-agent-trace-propagation",
        "markdown": "https://changegamer.ai/articles/multi-agent-trace-propagation.md",
        "json": "https://changegamer.ai/api/articles/multi-agent-trace-propagation.json"
      },
      {
        "slug": "llm-as-judge-screening-in-production",
        "title": "Designing an LLM-as-Judge Pipeline for Production AI Agents",
        "description": "An operator playbook for screening live AI agent output with an LLM judge: a confidence/stakes routing architecture to a human queue, continuous live-traffic rubric design, and per-bias mitigations for position, verbosity, and self-preference.",
        "kind": "sub",
        "order": 2,
        "html": "https://changegamer.ai/articles/llm-as-judge-screening-in-production",
        "markdown": "https://changegamer.ai/articles/llm-as-judge-screening-in-production.md",
        "json": "https://changegamer.ai/api/articles/llm-as-judge-screening-in-production.json"
      },
      {
        "slug": "agent-cost-telemetry-in-production",
        "title": "How to Track AI Agent Costs in Production",
        "description": "How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.",
        "kind": "sub",
        "order": 3,
        "html": "https://changegamer.ai/articles/agent-cost-telemetry-in-production",
        "markdown": "https://changegamer.ai/articles/agent-cost-telemetry-in-production.md",
        "json": "https://changegamer.ai/api/articles/agent-cost-telemetry-in-production.json"
      },
      {
        "slug": "incident-to-eval-fixture-loop",
        "title": "How to Turn an AI Agent Incident into an Evaluation Test Case",
        "description": "How to turn an AI agent incident into an evaluation test case: freeze the trace, redact it, minimize it to the failing decision, label the expected outcome with a trusted oracle, prove it fails then passes, and retire it later.",
        "kind": "sub",
        "order": 4,
        "html": "https://changegamer.ai/articles/incident-to-eval-fixture-loop",
        "markdown": "https://changegamer.ai/articles/incident-to-eval-fixture-loop.md",
        "json": "https://changegamer.ai/api/articles/incident-to-eval-fixture-loop.json"
      },
      {
        "slug": "retrieval-attribution-logging-in-production",
        "title": "How to Log RAG Retrieval in Production for Debugging Agent Answers",
        "description": "How to log RAG retrieval in production so a wrong agent answer can be debugged: per-retrieval fields, what reached the context window versus what the model cited, privacy by hashes and IDs, and retention and sampling choices.",
        "kind": "sub",
        "order": 5,
        "html": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production",
        "markdown": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production.md",
        "json": "https://changegamer.ai/api/articles/retrieval-attribution-logging-in-production.json"
      },
      {
        "slug": "online-feedback-signals-for-ai-agents",
        "title": "How to Measure AI Agent Quality from Live User Feedback",
        "description": "How to measure AI agent quality from live user feedback: why explicit ratings are sparse and biased, how re-asks, abandonment and escalations mislead, and how to join each signal to a trace and route it to review.",
        "kind": "sub",
        "order": 6,
        "html": "https://changegamer.ai/articles/online-feedback-signals-for-ai-agents",
        "markdown": "https://changegamer.ai/articles/online-feedback-signals-for-ai-agents.md",
        "json": "https://changegamer.ai/api/articles/online-feedback-signals-for-ai-agents.json"
      },
      {
        "slug": "online-agent-metrics-and-drift-monitoring",
        "title": "How to Detect Quality Drift in a Production AI Agent",
        "description": "How to detect quality drift in a production AI agent: baseline aggregate signals, alert on a diff against the baseline, and separate a prompt, model or tool-version change from a shift in traffic mix.",
        "kind": "sub",
        "order": 7,
        "html": "https://changegamer.ai/articles/online-agent-metrics-and-drift-monitoring",
        "markdown": "https://changegamer.ai/articles/online-agent-metrics-and-drift-monitoring.md",
        "json": "https://changegamer.ai/api/articles/online-agent-metrics-and-drift-monitoring.json"
      },
      {
        "slug": "sampling-and-retention-of-agent-traces",
        "title": "How to Sample and Retain Production AI Agent Traces",
        "description": "How to sample and retain production AI agent traces: head vs tail sampling, keep-all-errors plus a random baseline, whole-trace decisions for multi-agent runs, retention tiers and redaction before the clock starts.",
        "kind": "sub",
        "order": 8,
        "html": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces",
        "markdown": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/sampling-and-retention-of-agent-traces.json"
      },
      {
        "slug": "crash-safe-run-start-records-for-agent-traces",
        "title": "How to Keep Trace Data When an AI Agent Crashes Mid-Run",
        "description": "How to keep trace data when an AI agent crashes mid-run: write a small run-start record outside the trace buffer, detect orphans, count them as unknown outcomes, and force-decide on shutdown.",
        "kind": "sub",
        "order": 9,
        "html": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces",
        "markdown": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/crash-safe-run-start-records-for-agent-traces.json"
      },
      {
        "slug": "choosing-an-llm-observability-backend",
        "title": "How to Choose an LLM Observability Platform for AI Agents",
        "description": "How to choose an LLM observability platform: decide on OTel-native ingestion, self-host versus cloud, export portability, redaction hooks and retention support before comparing vendors.",
        "kind": "sub",
        "order": 10,
        "html": "https://changegamer.ai/articles/choosing-an-llm-observability-backend",
        "markdown": "https://changegamer.ai/articles/choosing-an-llm-observability-backend.md",
        "json": "https://changegamer.ai/api/articles/choosing-an-llm-observability-backend.json"
      },
      {
        "slug": "redacting-sensitive-data-from-agent-traces",
        "title": "How to Redact PII From AI Agent Traces: Placement, Testing and Cleanup",
        "description": "How to redact PII from AI agent traces in practice: where the redactor sits in the pipeline, what to do per span field, how to test it with seeded fake PII, and how to clean up after a leak.",
        "kind": "sub",
        "order": 11,
        "html": "https://changegamer.ai/articles/redacting-sensitive-data-from-agent-traces",
        "markdown": "https://changegamer.ai/articles/redacting-sensitive-data-from-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/redacting-sensitive-data-from-agent-traces.json"
      }
    ]
  },
  "navigation": {
    "pillar": {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "kind": "pillar",
      "order": 0,
      "html": "https://changegamer.ai/articles/agent-observability-and-evaluation",
      "markdown": "https://changegamer.ai/articles/agent-observability-and-evaluation.md",
      "json": "https://changegamer.ai/api/articles/agent-observability-and-evaluation.json"
    },
    "previous": {
      "slug": "multi-agent-trace-propagation",
      "title": "Distributed Tracing for Multi-Agent AI Systems",
      "description": "How trace ID propagation actually works across a multi-agent handoff, why handoff and delegation need different trace shapes, and why an orchestrator-level trace can hide a failed leg of a fan-out.",
      "kind": "sub",
      "order": 1,
      "html": "https://changegamer.ai/articles/multi-agent-trace-propagation",
      "markdown": "https://changegamer.ai/articles/multi-agent-trace-propagation.md",
      "json": "https://changegamer.ai/api/articles/multi-agent-trace-propagation.json"
    },
    "next": {
      "slug": "agent-cost-telemetry-in-production",
      "title": "How to Track AI Agent Costs in Production",
      "description": "How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.",
      "kind": "sub",
      "order": 3,
      "html": "https://changegamer.ai/articles/agent-cost-telemetry-in-production",
      "markdown": "https://changegamer.ai/articles/agent-cost-telemetry-in-production.md",
      "json": "https://changegamer.ai/api/articles/agent-cost-telemetry-in-production.json"
    }
  },
  "resources": [
    {
      "slug": "evaluating-ai-agents",
      "html": "https://changegamer.ai/resources/evaluating-ai-agents",
      "markdown": "https://changegamer.ai/resources/evaluating-ai-agents.md",
      "json": "https://changegamer.ai/api/resources/evaluating-ai-agents.json"
    },
    {
      "slug": "agent-observability",
      "html": "https://changegamer.ai/resources/agent-observability",
      "markdown": "https://changegamer.ai/resources/agent-observability.md",
      "json": "https://changegamer.ai/api/resources/agent-observability.json"
    }
  ]
}