{
  "slug": "agent-cost-telemetry-in-production",
  "title": "How to Track AI Agent Costs in Production",
  "description": "How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.",
  "kind": "sub",
  "order": 3,
  "target_query": "how to track AI agent costs in production",
  "secondary_queries": [
    "opentelemetry genai",
    "llm cost tracking",
    "ai agent cost attribution",
    "llm budget alerts"
  ],
  "tags": [
    "agents",
    "cost",
    "observability",
    "telemetry",
    "opentelemetry",
    "production"
  ],
  "published": "2026-09-29",
  "updated": "2026-09-29",
  "words": 1553,
  "estimated_tokens": 2065,
  "premium": false,
  "rights": {
    "access": "free",
    "note": "Editorial guides are always free and never part of the licensed corpus.",
    "license": "https://changegamer.ai/license.xml",
    "pricing": "https://changegamer.ai/api/pricing.json",
    "payment": "https://changegamer.ai/api/payment.json"
  },
  "license": "https://changegamer.ai/license.xml",
  "citation": "ChangeGamer (2026-09-29). How to Track AI Agent Costs in Production. ChangeGamer. https://changegamer.ai/articles/agent-cost-telemetry-in-production (updated 2026-09-29).",
  "bibtex": "@misc{changegamer_agent_cost_telemetry_in_production, title = {How to Track AI Agent Costs in Production}, publisher = {ChangeGamer}, year = {2026}, url = {https://changegamer.ai/articles/agent-cost-telemetry-in-production}, note = {Updated 2026-09-29}}",
  "canonical": "https://changegamer.ai/articles/agent-cost-telemetry-in-production",
  "markdown": "https://changegamer.ai/articles/agent-cost-telemetry-in-production.md",
  "takeaways": [
    "Tracking AI agent cost in production means storing raw token counts on every span and computing dollars later against a versioned price table, so a provider price change never rewrites the cost of past runs.",
    "Cached and uncached input tokens should be recorded as separate counts on each model-call span, because a single blended input-token number cannot reproduce the invoice once cache hits are billed at a different rate.",
    "Retried, failed and abandoned agent runs spend real money, so cost telemetry should attribute their spend to the task that caused it instead of dropping them from the denominator of cost per successful task.",
    "Trace sampling must never apply to cost accounting: a sampled span can be discarded for storage reasons, but its token counts still need to reach a cost counter, or the total will drift below the invoice.",
    "Budget alerts for agents work best in three layers: soft and hard thresholds on cumulative spend, a per-run kill limit that stops a runaway loop, and a rate-of-spend alert that fires before a cap is reached.",
    "ChangeGamer's scripts/seo-budget.mjs enforces a USD 5.00 monthly hard cap on third-party measurement API spend by pairing a pessimistic preflight authorize step with a record step that writes the provider-reported actual."
  ],
  "outline": [
    {
      "depth": 2,
      "text": "What should a span record so its cost can be recomputed later?",
      "anchor": "what-should-a-span-record-so-its-cost-can-be-recomputed-later",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#what-should-a-span-record-so-its-cost-can-be-recomputed-later"
    },
    {
      "depth": 2,
      "text": "Versioning the price table so historical cost stays stable",
      "anchor": "versioning-the-price-table-so-historical-cost-stays-stable",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#versioning-the-price-table-so-historical-cost-stays-stable"
    },
    {
      "depth": 2,
      "text": "Why record cached and uncached input tokens separately?",
      "anchor": "why-record-cached-and-uncached-input-tokens-separately",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#why-record-cached-and-uncached-input-tokens-separately"
    },
    {
      "depth": 2,
      "text": "Attributing retried and abandoned runs to the task that caused them",
      "anchor": "attributing-retried-and-abandoned-runs-to-the-task-that-caused-them",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#attributing-retried-and-abandoned-runs-to-the-task-that-caused-them"
    },
    {
      "depth": 2,
      "text": "Attributing shared costs to tenants",
      "anchor": "attributing-shared-costs-to-tenants",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#attributing-shared-costs-to-tenants"
    },
    {
      "depth": 2,
      "text": "Keeping sampling out of cost accounting",
      "anchor": "keeping-sampling-out-of-cost-accounting",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#keeping-sampling-out-of-cost-accounting"
    },
    {
      "depth": 2,
      "text": "How do you reconcile estimated cost against the provider invoice?",
      "anchor": "how-do-you-reconcile-estimated-cost-against-the-provider-invoice",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#how-do-you-reconcile-estimated-cost-against-the-provider-invoice"
    },
    {
      "depth": 2,
      "text": "Budget guards for AI agents: thresholds, kill limits and rate of spend",
      "anchor": "budget-guards-for-ai-agents-thresholds-kill-limits-and-rate-of-spend",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#budget-guards-for-ai-agents-thresholds-kill-limits-and-rate-of-spend"
    },
    {
      "depth": 3,
      "text": "A first-hand example: authorize, then record",
      "anchor": "a-first-hand-example-authorize-then-record",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#a-first-hand-example-authorize-then-record"
    },
    {
      "depth": 2,
      "text": "Where do the cost-reduction levers live?",
      "anchor": "where-do-the-cost-reduction-levers-live",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#where-do-the-cost-reduction-levers-live"
    },
    {
      "depth": 2,
      "text": "Sources and further reading",
      "anchor": "sources-and-further-reading",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production#sources-and-further-reading"
    }
  ],
  "faq": [
    {
      "question": "Should an AI agent store cost in dollars or in token counts?",
      "answer": "An AI agent should store token counts, split into input, cached input and output, on each span, and compute dollars at query time from a dated price table. Stored dollar figures go stale when a provider changes its rates and cannot be recomputed, while raw token counts can always be repriced."
    },
    {
      "question": "Why does my AI agent cost dashboard not match the provider invoice?",
      "answer": "A dashboard usually drifts from the invoice because of an out-of-date price table, cached tokens billed at a different rate than the dashboard assumes, spans dropped by sampling, or spend that never passed through the instrumented code path. Reconciling a dashboard against the invoice on a fixed cadence exposes which of these it is."
    },
    {
      "question": "How do I attribute shared costs like a system prompt to individual tenants?",
      "answer": "Shared costs such as a system prompt or a retrieval index have no single owner, so teams must choose an allocation rule, for example splitting by each tenant's share of requests, and apply it consistently. No published standard prescribes one rule, so the choice is a policy decision worth documenting."
    },
    {
      "question": "Does OpenTelemetry cover AI agent cost tracking?",
      "answer": "OpenTelemetry's GenAI conventions cover the inputs to cost tracking, such as the model name and gen_ai.usage.input_tokens and gen_ai.usage.output_tokens on a span, but the price lookup and the dollar figure are computed by your own code or your observability backend. As of September 2026 the GenAI attribute vocabulary is still labeled Development rather than stable."
    }
  ],
  "body": "Cost tracking for a production AI agent is a bookkeeping problem before it is a monitoring problem: the question is whether the number you report can be reproduced from stored facts, and whether it will still match the invoice next month. [The pillar](/articles/agent-observability-and-evaluation) defines cost telemetry and its aggregation axes, and [agent observability for reliability](/articles/agent-observability-for-reliability) covers the per-span cost field; this article covers what breaks when that field meets real billing. The mechanisms below are this article's own reasoning, not a published standard.\n\n## What should a span record so its cost can be recomputed later?\n\nA span should record raw usage, not a finished dollar figure, so cost can be recomputed whenever the price changes. The agent-observability resource lists `gen_ai.request.model`, `gen_ai.usage.input_tokens` and `gen_ai.usage.output_tokens` among the OpenTelemetry GenAI attributes, and says the GenAI vocabulary is still labeled Development rather than final; as of September 2026, treat those names as likely to shift and keep a thin mapping layer of your own.\n\nStore three counts per model call: uncached input tokens, cached input tokens, and output tokens. Then derive dollars at query time. If you store only a computed dollar amount, every later price change leaves you unable to tell whether a cost movement came from behavior or from a rate card.\n\n## Versioning the price table so historical cost stays stable\n\nA price table keeps historical cost stable when it is versioned with an effective date and each span is priced against the version in force when it ran. A provider that lowers a rate would otherwise make last quarter's runs look cheaper after a recompute, and a raised rate would make them look worse, so a cost trend chart would move without any change in agent behavior.\n\nA workable layout:\n\n- One row per model and price component, each with a start date and an end date.\n- The span stores the model identifier and token counts; the cost query joins on the span's timestamp.\n\nChangeGamer's own `scripts/seo-budget.mjs` shows the failure mode on a small scale. Each price in its table carries a `verified` date and a `source`, and prices marked `estimated` get a safety multiplier of 2. Its comment on the `aeo.llm_response` entry, verified 2026-09-04, records that an earlier `$0.01` figure was an 8x under-estimate, corrected after measured calls came back at roughly `$0.072` and `$0.02`. A stale price was invisible until real actuals were compared to it.\n\n## Why record cached and uncached input tokens separately?\n\nCached and uncached input tokens are billed differently by the provider, so a single blended input count cannot reproduce what you will be charged. The [agent cost and latency optimization resource](/resources/agent-cost-latency-optimization) states that every major provider applies some discount on cache hits but that the percentage and minimum prompt size differ per vendor. That vendor variance is a telemetry problem in its own right: a fixed discount assumed in one place will be wrong for at least one provider.\n\nRecording the split also makes cache behavior observable. A falling cached share of input tokens on an unchanged prompt is a cost regression you can only see if the two counts were never merged.\n\n## Attributing retried and abandoned runs to the task that caused them\n\nRetried and abandoned runs should be charged to the task that triggered them, and their cost must stay in the books even when the user never saw an answer. Retries, tool-call failures and runs the user abandoned all consumed tokens, so dropping them makes cost per task look better than the invoice supports.\n\nTrack three numbers instead of one:\n\n| Metric | What it includes | What it answers |\n|---|---|---|\n| Cost per attempt | Every run, including retries | What did we spend? |\n| Cost per completed task | All attempts' spend, divided by completed tasks | What does a delivered result cost? |\n| Waste share | Spend on failed, retried or abandoned attempts | How much of the bill bought nothing? |\n\nDefine \"abandoned\" explicitly, because the definition changes the waste share. Attaching a retry to its parent run's trace, using the trace ID the observability resource describes, is what lets a retry's spend roll up to the original task rather than appearing as an unrelated request.\n\n## Attributing shared costs to tenants\n\nShared costs have no natural owner, so you must pick an allocation rule and apply it consistently. A long system prompt read by every tenant's requests, a retrieval index, or a background summarization job all serve many tenants at once. No corpus resource or published standard prescribes an allocation rule, so treat the following as options, not a recommendation:\n\n- **By request share**: each tenant carries its fraction of total requests. Simple, but it under-charges a tenant whose requests are unusually large.\n- **By token share**: allocate in proportion to tokens consumed. It tracks usage more closely, and it needs the token counts to be trustworthy.\n- **Unallocated overhead**: keep shared spend as its own line and never distribute it. It avoids arbitrary splits, at the price of an unattributed remainder.\n\nTag each tenant's own model calls at the span level; shared costs are the residue. Documenting the rule matters more than which one you pick.\n\n## Keeping sampling out of cost accounting\n\nSampling is acceptable for stored spans but never for cost accounting, because a discarded span's tokens were still billed. Even a team that samples its stored traces should increment a cost counter from every model call before any sampling decision is made.\n\nThe practical shape is two paths from the same instrumentation point. One path emits a span that a sampler may drop. The other adds token counts to a metric counter, keyed by model and by whatever tenant or task labels you have, that no sampler touches. Keep counter labels low in cardinality and put per-run detail in the span.\n\n## How do you reconcile estimated cost against the provider invoice?\n\nReconcile by comparing your computed spend to the provider's billed total on a fixed cadence and treating the difference as a signal to investigate, not a rounding error. Common causes are a stale price row, an unrecorded cached-token rate, sampled-out spans, and calls made outside the instrumented path, such as a script or a notebook using the same API key.\n\nChangeGamer's spend governor makes this discipline explicit. Its ledger in `agents/SEO-AEO-LEDGER.md` states that actuals come from the provider, and that when a call returns no cost field the row is recorded as an estimate and labeled `estimate` in the `Source` column, so it is visibly weaker than its neighbors. If the ledger and the provider's billing page disagree, the provider is right and the file gets a correcting row noted as a reconciliation. For agent telemetry, mark every cost figure as measured or estimated, and let the invoice win.\n\n## Budget guards for AI agents: thresholds, kill limits and rate of spend\n\nA budget guard has three layers: cumulative thresholds, a per-run limit, and a rate-of-spend alert. Each catches a failure the others miss.\n\n1. **Soft and hard thresholds** on cumulative spend per period. Crossing the soft one notifies a human; crossing the hard one stops discretionary calls.\n2. **A per-run kill limit** that aborts a single trace once its running cost passes a ceiling, which is the guard against a loop that never terminates. Fan-out makes this necessary, as the cost resource notes that sub-agent dispatch multiplies spend well beyond a single-call baseline.\n3. **A rate-of-spend alert** comparing spend over the last hour to a normal hour, which fires days before a monthly threshold would.\n\nNo corpus resource specifies threshold values, and this article does not invent them; derive yours from your own baseline.\n\n### A first-hand example: authorize, then record\n\nChangeGamer's `scripts/seo-budget.mjs` runs exactly a soft/hard threshold guard with a preflight and a ledger. To be clear about its scope, it governs third-party SEO and measurement API spend against a USD 5.00 monthly hard cap. It does not meter LLM tokens.\n\n```bash\nnode scripts/seo-budget.mjs status\nnode scripts/seo-budget.mjs authorize --item aeo.llm_response --qty 4\n# exit 0 = authorized, exit 1 = denied\nnode scripts/seo-budget.mjs record --sku aeo.llm_response --qty 4 --cost <provider-reported-usd> --cycle <id>\n```\n\n`authorize` adds the estimate for the requested calls to month-to-date spend and exits 1 if the projection would cross the USD 5.00 hard cap, the USD 3.80 soft cap, or a per-category sub-budget. Prices marked `estimated` are multiplied by 2 in that projection. `record` then writes the provider-reported actual, and if the ledger ends up over the cap it still writes the row but exits 1, because the ledger must stay honest. Estimates gate the call; actuals govern the books. That split transfers directly to agents: a pessimistic preflight before a costly run, measured usage after it.\n\n## Where do the cost-reduction levers live?\n\nThe levers for reducing agent spend live in the [agent cost and latency optimization resource](/resources/agent-cost-latency-optimization), and this article deliberately does not list them. Telemetry shows where the money went; choosing what to change is a separate discipline. For the trace-level view these numbers roll up into, return to [the pillar](/articles/agent-observability-and-evaluation).\n\n## Sources and further reading\n\nToken-usage attributes and the Development status of the GenAI conventions are in [agent observability and tracing](/resources/agent-observability). The spend-governor behavior is read directly from ChangeGamer's `scripts/seo-budget.mjs` and `agents/SEO-AEO-LEDGER.md` as of 29 September 2026.",
  "cluster": {
    "id": "agent-observability-evaluation",
    "title": "Agent observability and evaluation",
    "description": "How to observe and evaluate an AI agent already live in production — tracing spans for tool calls and retrieval steps, judge-based screening versus human acceptance of live output, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures — not the pre-ship, CI-gating evaluation methodology already owned by evaluating-ai-agents-in-ci, and not the trace/span field mechanics already owned by agent-observability-for-reliability, both in the completed agent-reliability cluster.",
    "status": "complete",
    "pillar": {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "kind": "pillar",
      "order": 0,
      "html": "https://changegamer.ai/articles/agent-observability-and-evaluation",
      "markdown": "https://changegamer.ai/articles/agent-observability-and-evaluation.md",
      "json": "https://changegamer.ai/api/articles/agent-observability-and-evaluation.json"
    },
    "articles": [
      {
        "slug": "multi-agent-trace-propagation",
        "title": "Distributed Tracing for Multi-Agent AI Systems",
        "description": "How trace ID propagation actually works across a multi-agent handoff, why handoff and delegation need different trace shapes, and why an orchestrator-level trace can hide a failed leg of a fan-out.",
        "kind": "sub",
        "order": 1,
        "html": "https://changegamer.ai/articles/multi-agent-trace-propagation",
        "markdown": "https://changegamer.ai/articles/multi-agent-trace-propagation.md",
        "json": "https://changegamer.ai/api/articles/multi-agent-trace-propagation.json"
      },
      {
        "slug": "llm-as-judge-screening-in-production",
        "title": "Designing an LLM-as-Judge Pipeline for Production AI Agents",
        "description": "An operator playbook for screening live AI agent output with an LLM judge: a confidence/stakes routing architecture to a human queue, continuous live-traffic rubric design, and per-bias mitigations for position, verbosity, and self-preference.",
        "kind": "sub",
        "order": 2,
        "html": "https://changegamer.ai/articles/llm-as-judge-screening-in-production",
        "markdown": "https://changegamer.ai/articles/llm-as-judge-screening-in-production.md",
        "json": "https://changegamer.ai/api/articles/llm-as-judge-screening-in-production.json"
      },
      {
        "slug": "agent-cost-telemetry-in-production",
        "title": "How to Track AI Agent Costs in Production",
        "description": "How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.",
        "kind": "sub",
        "order": 3,
        "html": "https://changegamer.ai/articles/agent-cost-telemetry-in-production",
        "markdown": "https://changegamer.ai/articles/agent-cost-telemetry-in-production.md",
        "json": "https://changegamer.ai/api/articles/agent-cost-telemetry-in-production.json"
      },
      {
        "slug": "incident-to-eval-fixture-loop",
        "title": "How to Turn an AI Agent Incident into an Evaluation Test Case",
        "description": "How to turn an AI agent incident into an evaluation test case: freeze the trace, redact it, minimize it to the failing decision, label the expected outcome with a trusted oracle, prove it fails then passes, and retire it later.",
        "kind": "sub",
        "order": 4,
        "html": "https://changegamer.ai/articles/incident-to-eval-fixture-loop",
        "markdown": "https://changegamer.ai/articles/incident-to-eval-fixture-loop.md",
        "json": "https://changegamer.ai/api/articles/incident-to-eval-fixture-loop.json"
      },
      {
        "slug": "retrieval-attribution-logging-in-production",
        "title": "How to Log RAG Retrieval in Production for Debugging Agent Answers",
        "description": "How to log RAG retrieval in production so a wrong agent answer can be debugged: per-retrieval fields, what reached the context window versus what the model cited, privacy by hashes and IDs, and retention and sampling choices.",
        "kind": "sub",
        "order": 5,
        "html": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production",
        "markdown": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production.md",
        "json": "https://changegamer.ai/api/articles/retrieval-attribution-logging-in-production.json"
      },
      {
        "slug": "online-feedback-signals-for-ai-agents",
        "title": "How to Measure AI Agent Quality from Live User Feedback",
        "description": "How to measure AI agent quality from live user feedback: why explicit ratings are sparse and biased, how re-asks, abandonment and escalations mislead, and how to join each signal to a trace and route it to review.",
        "kind": "sub",
        "order": 6,
        "html": "https://changegamer.ai/articles/online-feedback-signals-for-ai-agents",
        "markdown": "https://changegamer.ai/articles/online-feedback-signals-for-ai-agents.md",
        "json": "https://changegamer.ai/api/articles/online-feedback-signals-for-ai-agents.json"
      },
      {
        "slug": "online-agent-metrics-and-drift-monitoring",
        "title": "How to Detect Quality Drift in a Production AI Agent",
        "description": "How to detect quality drift in a production AI agent: baseline aggregate signals, alert on a diff against the baseline, and separate a prompt, model or tool-version change from a shift in traffic mix.",
        "kind": "sub",
        "order": 7,
        "html": "https://changegamer.ai/articles/online-agent-metrics-and-drift-monitoring",
        "markdown": "https://changegamer.ai/articles/online-agent-metrics-and-drift-monitoring.md",
        "json": "https://changegamer.ai/api/articles/online-agent-metrics-and-drift-monitoring.json"
      },
      {
        "slug": "sampling-and-retention-of-agent-traces",
        "title": "How to Sample and Retain Production AI Agent Traces",
        "description": "How to sample and retain production AI agent traces: head vs tail sampling, keep-all-errors plus a random baseline, whole-trace decisions for multi-agent runs, retention tiers and redaction before the clock starts.",
        "kind": "sub",
        "order": 8,
        "html": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces",
        "markdown": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/sampling-and-retention-of-agent-traces.json"
      },
      {
        "slug": "crash-safe-run-start-records-for-agent-traces",
        "title": "How to Keep Trace Data When an AI Agent Crashes Mid-Run",
        "description": "How to keep trace data when an AI agent crashes mid-run: write a small run-start record outside the trace buffer, detect orphans, count them as unknown outcomes, and force-decide on shutdown.",
        "kind": "sub",
        "order": 9,
        "html": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces",
        "markdown": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/crash-safe-run-start-records-for-agent-traces.json"
      },
      {
        "slug": "choosing-an-llm-observability-backend",
        "title": "How to Choose an LLM Observability Platform for AI Agents",
        "description": "How to choose an LLM observability platform: decide on OTel-native ingestion, self-host versus cloud, export portability, redaction hooks and retention support before comparing vendors.",
        "kind": "sub",
        "order": 10,
        "html": "https://changegamer.ai/articles/choosing-an-llm-observability-backend",
        "markdown": "https://changegamer.ai/articles/choosing-an-llm-observability-backend.md",
        "json": "https://changegamer.ai/api/articles/choosing-an-llm-observability-backend.json"
      },
      {
        "slug": "redacting-sensitive-data-from-agent-traces",
        "title": "How to Redact PII From AI Agent Traces: Placement, Testing and Cleanup",
        "description": "How to redact PII from AI agent traces in practice: where the redactor sits in the pipeline, what to do per span field, how to test it with seeded fake PII, and how to clean up after a leak.",
        "kind": "sub",
        "order": 11,
        "html": "https://changegamer.ai/articles/redacting-sensitive-data-from-agent-traces",
        "markdown": "https://changegamer.ai/articles/redacting-sensitive-data-from-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/redacting-sensitive-data-from-agent-traces.json"
      }
    ]
  },
  "navigation": {
    "pillar": {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "kind": "pillar",
      "order": 0,
      "html": "https://changegamer.ai/articles/agent-observability-and-evaluation",
      "markdown": "https://changegamer.ai/articles/agent-observability-and-evaluation.md",
      "json": "https://changegamer.ai/api/articles/agent-observability-and-evaluation.json"
    },
    "previous": {
      "slug": "llm-as-judge-screening-in-production",
      "title": "Designing an LLM-as-Judge Pipeline for Production AI Agents",
      "description": "An operator playbook for screening live AI agent output with an LLM judge: a confidence/stakes routing architecture to a human queue, continuous live-traffic rubric design, and per-bias mitigations for position, verbosity, and self-preference.",
      "kind": "sub",
      "order": 2,
      "html": "https://changegamer.ai/articles/llm-as-judge-screening-in-production",
      "markdown": "https://changegamer.ai/articles/llm-as-judge-screening-in-production.md",
      "json": "https://changegamer.ai/api/articles/llm-as-judge-screening-in-production.json"
    },
    "next": {
      "slug": "incident-to-eval-fixture-loop",
      "title": "How to Turn an AI Agent Incident into an Evaluation Test Case",
      "description": "How to turn an AI agent incident into an evaluation test case: freeze the trace, redact it, minimize it to the failing decision, label the expected outcome with a trusted oracle, prove it fails then passes, and retire it later.",
      "kind": "sub",
      "order": 4,
      "html": "https://changegamer.ai/articles/incident-to-eval-fixture-loop",
      "markdown": "https://changegamer.ai/articles/incident-to-eval-fixture-loop.md",
      "json": "https://changegamer.ai/api/articles/incident-to-eval-fixture-loop.json"
    }
  },
  "resources": [
    {
      "slug": "agent-observability",
      "html": "https://changegamer.ai/resources/agent-observability",
      "markdown": "https://changegamer.ai/resources/agent-observability.md",
      "json": "https://changegamer.ai/api/resources/agent-observability.json"
    },
    {
      "slug": "agent-cost-latency-optimization",
      "html": "https://changegamer.ai/resources/agent-cost-latency-optimization",
      "markdown": "https://changegamer.ai/resources/agent-cost-latency-optimization.md",
      "json": "https://changegamer.ai/api/resources/agent-cost-latency-optimization.json"
    }
  ]
}