{
  "slug": "incident-to-eval-fixture-loop",
  "title": "How to Turn an AI Agent Incident into an Evaluation Test Case",
  "description": "How to turn an AI agent incident into an evaluation test case: freeze the trace, redact it, minimize it to the failing decision, label the expected outcome with a trusted oracle, prove it fails then passes, and retire it later.",
  "kind": "sub",
  "order": 4,
  "target_query": "how to turn an AI agent incident into an evaluation test case",
  "secondary_queries": [
    "llm evaluation dataset",
    "ai agent regression testing",
    "regression test from production failure",
    "agent eval fixture"
  ],
  "tags": [
    "agents",
    "evaluation",
    "incidents",
    "testing",
    "fixtures",
    "production"
  ],
  "published": "2026-09-30",
  "updated": "2026-09-30",
  "words": 1537,
  "estimated_tokens": 2044,
  "premium": false,
  "rights": {
    "access": "free",
    "note": "Editorial guides are always free and never part of the licensed corpus.",
    "license": "https://changegamer.ai/license.xml",
    "pricing": "https://changegamer.ai/api/pricing.json",
    "payment": "https://changegamer.ai/api/payment.json"
  },
  "license": "https://changegamer.ai/license.xml",
  "citation": "ChangeGamer (2026-09-30). How to Turn an AI Agent Incident into an Evaluation Test Case. ChangeGamer. https://changegamer.ai/articles/incident-to-eval-fixture-loop (updated 2026-09-30).",
  "bibtex": "@misc{changegamer_incident_to_eval_fixture_loop, title = {How to Turn an AI Agent Incident into an Evaluation Test Case}, publisher = {ChangeGamer}, year = {2026}, url = {https://changegamer.ai/articles/incident-to-eval-fixture-loop}, note = {Updated 2026-09-30}}",
  "canonical": "https://changegamer.ai/articles/incident-to-eval-fixture-loop",
  "markdown": "https://changegamer.ai/articles/incident-to-eval-fixture-loop.md",
  "takeaways": [
    "Turning an AI agent incident into an evaluation test case starts with freezing the production trace before any rollback, because recovery steps can destroy the only faithful record of the failing run.",
    "A trace must be redacted of personal data and secrets before it becomes a committed test fixture, since a fixture lives in version control far longer and is read by far more people than the original trace.",
    "An incident fixture should be minimized to the single failing decision, keeping only the context that changes the outcome, so the test stays readable and does not break every time an unrelated prompt detail changes.",
    "The expected outcome of an incident fixture should be labeled by a human or a deterministic check, not by an LLM judge alone, because a judge cannot certify a case it may itself have misjudged in production.",
    "A new incident fixture is only trustworthy once it has failed against the buggy version and then passed against the fix, the same red-then-green discipline used for any regression test.",
    "Some incidents deserve a machine-checked fixture while others are better served by a written checklist rule, and ChangeGamer's July 2026 missed-windows incident became a protocol rule plus a status-board detector rather than a build check."
  ],
  "outline": [
    {
      "depth": 2,
      "text": "What has to happen before you roll back?",
      "anchor": "what-has-to-happen-before-you-roll-back",
      "url": "https://changegamer.ai/articles/incident-to-eval-fixture-loop#what-has-to-happen-before-you-roll-back"
    },
    {
      "depth": 2,
      "text": "Why redact before the trace becomes a fixture?",
      "anchor": "why-redact-before-the-trace-becomes-a-fixture",
      "url": "https://changegamer.ai/articles/incident-to-eval-fixture-loop#why-redact-before-the-trace-becomes-a-fixture"
    },
    {
      "depth": 2,
      "text": "How do you minimize a trace to the failing decision?",
      "anchor": "how-do-you-minimize-a-trace-to-the-failing-decision",
      "url": "https://changegamer.ai/articles/incident-to-eval-fixture-loop#how-do-you-minimize-a-trace-to-the-failing-decision"
    },
    {
      "depth": 2,
      "text": "How should you label the expected outcome?",
      "anchor": "how-should-you-label-the-expected-outcome",
      "url": "https://changegamer.ai/articles/incident-to-eval-fixture-loop#how-should-you-label-the-expected-outcome"
    },
    {
      "depth": 2,
      "text": "How do you prove the fixture is red, then green?",
      "anchor": "how-do-you-prove-the-fixture-is-red-then-green",
      "url": "https://changegamer.ai/articles/incident-to-eval-fixture-loop#how-do-you-prove-the-fixture-is-red-then-green"
    },
    {
      "depth": 2,
      "text": "Which layer should the fixture run in?",
      "anchor": "which-layer-should-the-fixture-run-in",
      "url": "https://changegamer.ai/articles/incident-to-eval-fixture-loop#which-layer-should-the-fixture-run-in"
    },
    {
      "depth": 2,
      "text": "How do you avoid an over-fitted fixture?",
      "anchor": "how-do-you-avoid-an-over-fitted-fixture",
      "url": "https://changegamer.ai/articles/incident-to-eval-fixture-loop#how-do-you-avoid-an-over-fitted-fixture"
    },
    {
      "depth": 2,
      "text": "What provenance should each fixture carry?",
      "anchor": "what-provenance-should-each-fixture-carry",
      "url": "https://changegamer.ai/articles/incident-to-eval-fixture-loop#what-provenance-should-each-fixture-carry"
    },
    {
      "depth": 2,
      "text": "When should you retire an incident fixture?",
      "anchor": "when-should-you-retire-an-incident-fixture",
      "url": "https://changegamer.ai/articles/incident-to-eval-fixture-loop#when-should-you-retire-an-incident-fixture"
    },
    {
      "depth": 2,
      "text": "Which incidents deserve a fixture, and which only a rule?",
      "anchor": "which-incidents-deserve-a-fixture-and-which-only-a-rule",
      "url": "https://changegamer.ai/articles/incident-to-eval-fixture-loop#which-incidents-deserve-a-fixture-and-which-only-a-rule"
    }
  ],
  "faq": [
    {
      "question": "What is an incident fixture in AI agent evaluation?",
      "answer": "An incident fixture is a labeled test case built from a real production failure: the recorded input or trajectory that went wrong, paired with the correct expected outcome, added to the suite that gates future releases. It turns one incident into a permanent regression check instead of a postmortem nobody re-reads."
    },
    {
      "question": "Can I use a raw production trace as a test case?",
      "answer": "A raw production trace should not be committed as a test case without redaction and minimization. Traces can contain personal data, secrets and unrelated context, and the testing-ai-agents resource warns to scrub credentials and sensitive headers from cassettes before committing them to version control."
    },
    {
      "question": "How do I know an incident test case actually catches the bug?",
      "answer": "Run the new test against the version that caused the incident and confirm it fails, then run it against the fixed version and confirm it passes. A fixture that never failed proves nothing, because it may be asserting something the buggy version already satisfied."
    },
    {
      "question": "Should every AI agent incident become an automated test?",
      "answer": "No. An incident whose cause is a repeatable, checkable behavior is a good fixture candidate, while an incident caused by a process gap, such as a missed handoff, may be better fixed with a written rule. Choosing per incident avoids a suite full of tests that guard nothing real."
    }
  ],
  "body": "Turning an incident into an evaluation test case is a sequence of small, deliberate conversions, each of which can fail quietly if skipped. [The pillar](/articles/agent-observability-and-evaluation) explains why a postmortem should feed the suite that gates releases, and [how to evaluate AI agents in CI](/articles/evaluating-ai-agents-in-ci) covers the suite itself; this article covers the craft between them. The steps below are this article's own reasoning as of 30 September 2026, not a published standard.\n\n## What has to happen before you roll back?\n\nFreeze the trace before you roll back, because the rollback can overwrite or expire exactly the evidence you need. Export the full trace tree, the model and prompt versions in force, and the tool outputs the run saw. [The incident response runbooks article](/articles/agent-incident-response-runbooks) treats freezing as an incident step; here it doubles as fixture raw material, because a fixture built from memory of what went wrong is a guess.\n\nThe [agent-observability resource](/resources/agent-observability) is the reason this is possible at all. It describes a trace as one complete agent run under a stable `trace_id`, with each LLM call, tool call and retrieval as a span. A frozen trace of that shape holds the input, the decision points and the tool results in one artifact.\n\n## Why redact before the trace becomes a fixture?\n\nRedact first because a fixture is committed to version control, where it outlives the trace store's retention and reaches far more readers. The agent-observability resource says to redact PII before logging tool-call inputs and outputs, and notes prompt and completion bodies are off by default in OpenTelemetry for PII safety. A trace exported during an incident may still carry values that a later, stricter pipeline would have removed.\n\nTreat the frozen trace as sensitive and the fixture as the sanitized derivative:\n\n- Replace personal data with stable placeholders so the case keeps its shape.\n- Remove credentials and sensitive headers. The [testing-ai-agents resource](/resources/testing-ai-agents) says to scrub credentials and sensitive headers from cassettes before committing, and names `filter_headers` and `filter_post_data_parameters` in VCR.py and pytest-recording for it.\n- Keep the unredacted original in restricted storage only if your retention rules allow it. No corpus resource sets a window, so that decision is yours.\n\n## How do you minimize a trace to the failing decision?\n\nMinimize by cutting the trace down to the smallest input that still reproduces the wrong decision, then confirming the failure survives each cut. An incident trace is often dozens of steps, but the defect usually lives in one: a tool argument built wrongly, a result misread, a retry that repeated a side effect.\n\nA workable order:\n\n1. Identify the span where the run first went wrong, not where the damage surfaced.\n2. Keep that span's input, the tool results that fed it, and the instructions that governed it.\n3. Delete everything else, re-running after each cut. If the failure disappears, restore what you removed.\n\nMinimizing also protects the suite. A fixture stuffed with full conversation history breaks whenever an unrelated prompt line changes, which teaches the team to ignore red builds.\n\n## How should you label the expected outcome?\n\nLabel the expected outcome with a human decision or a deterministic check, not an LLM judge alone. The incident happened because something already got a case wrong in production, and if a judge was part of that screening path, reusing it as the sole oracle risks encoding the same blind spot into the test.\n\nThe [evaluating-ai-agents resource](/resources/evaluating-ai-agents) lists position, verbosity and self-preference biases for LLM judges and says tool-call correctness is better measured by direct comparison against a reference answer. For an incident fixture that suggests a hierarchy:\n\n| Oracle | Use it when | Weakness |\n|---|---|---|\n| Deterministic assertion on a tool call or field | The right behavior is exact, such as the correct argument or no duplicate write | Cannot judge free text |\n| Human-written expected outcome | The right behavior needs judgment once, at authoring time | Costs reviewer time per case |\n| LLM judge with a human-verified label | The outcome is prose and the rubric is stable | Judge can drift; label must come from a person |\n\n[The judge-design article](/articles/llm-as-judge-screening-in-production) covers building a judge; this article only says to keep a person or an exact check as the final word on a fixture.\n\n## How do you prove the fixture is red, then green?\n\nProve it by running the new case against the buggy version, watching it fail, then running it against the fix and watching it pass. A fixture that passes on the buggy code is asserting something the bug never violated, so it adds cost without adding protection.\n\n```bash\n# on the fix branch: take the buggy code, keep the new fixture, expect a failure\ngit checkout <pre-fix-commit> -- src/\npytest tests/agents -k incident_1234    # must FAIL\ngit checkout HEAD -- src/\npytest tests/agents -k incident_1234    # must PASS\n```\n\nThe path and test name are placeholders for your own layout. What matters is the order: red on the old behavior, green on the new. Record both results in the pull request that adds the fixture.\n\n## Which layer should the fixture run in?\n\nChoose the cheapest layer that can still fail on the defect: a unit test if the bug is in code around the model, a cassette replay if it depends on a recorded model exchange, and a live run only if nothing else reproduces it. The testing-ai-agents resource describes three layers: mocked-model unit tests, cassette replay on every commit, and a small live set in a nightly or pre-release job. It says most bugs live in the glue code, not the model.\n\nThat gives a decision rule for incidents:\n\n- **Argument built wrongly or output parsed wrongly:** unit test with the model mocked.\n- **Behavior tied to one recorded model response:** cassette replay. The resource says CI should be configured so a missing cassette fails instead of silently making a live call, and its cross-links note production traces can seed new cassettes and test cases.\n- **Behavior that only appears with the live model:** a curated live scenario, kept few because the resource calls these runs costly and flaky.\n\nWhether every trace converts cleanly to a cassette depends on your tooling; the resource states the seeding idea, not a conversion procedure.\n\n## How do you avoid an over-fitted fixture?\n\nAvoid over-fitting by asserting the behavior that was wrong, not the exact text or path the buggy run produced. The testing-ai-agents resource advises asserting on structure and argument values rather than free-text reasoning, and asserting against a schema and key fields rather than exact prose, which tolerates benign rephrasing while catching real regressions.\n\nTwo failure modes recur. A fixture pinned to a full tool-call sequence fails whenever the agent legitimately takes a different route to the same correct result. A fixture pinned to wording fails on harmless rephrasing. Write the assertion as the invariant the incident violated, for example \"the refund tool is called at most once for this order\", not \"the agent's third message says X\".\n\n## What provenance should each fixture carry?\n\nEach fixture should carry the incident ID, the date, the affected versions and a one-line statement of the violated invariant, so a future reader can judge whether it still matters. A case named `test_case_47` becomes impossible to retire or trust.\n\nPut the provenance next to the case, in a comment or metadata file, and link the postmortem. This is the same habit of dating claims that keeps documentation honest, applied to tests.\n\n## When should you retire an incident fixture?\n\nRetire a fixture when the behavior it guards no longer exists, for example the tool was removed or the workflow redesigned, and record why in the same change. No corpus resource sets a retention rule for fixtures, so treat retirement as a review question, not a schedule. A suite that only grows accumulates cases that pass for reasons unrelated to their intent, and provenance is what makes the retirement decision cheap.\n\n## Which incidents deserve a fixture, and which only a rule?\n\nAn incident deserves a machine-checked fixture when its cause is a repeatable behavior a test can observe, and only a written rule when the cause is a process gap. ChangeGamer has one example of each, and both are its own history.\n\nThe 2026-07-27 and 2026-07-28 missed article windows, recorded in `agents/SEO-PLAN.md` under \"Incident: three missed windows\", produced no branch, PR or journal entry in three scheduled runs. Its root cause is unconfirmed; the leading hypothesis is that the first unit of work was too large. What came out of it was a written protocol rule, checkpoint discipline in section 3a of `.claude/commands/article-cycle.md` (push the branch before writing, resume an existing branch), plus a status-board check that flagged the misses. It is not a build- or CI-enforced invariant.\n\nBy contrast, `src/lib/articles.ts` fails the build when a sub-article does not link to its pillar. That rule is machine-checked, so a regression breaks the build instead of waiting for a person to notice.\n\nNeither form is better in general. The test is whether a machine can observe the failure. If it can, write the fixture. If it cannot, write the rule and, where possible, a detector.",
  "cluster": {
    "id": "agent-observability-evaluation",
    "title": "Agent observability and evaluation",
    "description": "How to observe and evaluate an AI agent already live in production — tracing spans for tool calls and retrieval steps, judge-based screening versus human acceptance of live output, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures — not the pre-ship, CI-gating evaluation methodology already owned by evaluating-ai-agents-in-ci, and not the trace/span field mechanics already owned by agent-observability-for-reliability, both in the completed agent-reliability cluster.",
    "status": "complete",
    "pillar": {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "kind": "pillar",
      "order": 0,
      "html": "https://changegamer.ai/articles/agent-observability-and-evaluation",
      "markdown": "https://changegamer.ai/articles/agent-observability-and-evaluation.md",
      "json": "https://changegamer.ai/api/articles/agent-observability-and-evaluation.json"
    },
    "articles": [
      {
        "slug": "multi-agent-trace-propagation",
        "title": "Distributed Tracing for Multi-Agent AI Systems",
        "description": "How trace ID propagation actually works across a multi-agent handoff, why handoff and delegation need different trace shapes, and why an orchestrator-level trace can hide a failed leg of a fan-out.",
        "kind": "sub",
        "order": 1,
        "html": "https://changegamer.ai/articles/multi-agent-trace-propagation",
        "markdown": "https://changegamer.ai/articles/multi-agent-trace-propagation.md",
        "json": "https://changegamer.ai/api/articles/multi-agent-trace-propagation.json"
      },
      {
        "slug": "llm-as-judge-screening-in-production",
        "title": "Designing an LLM-as-Judge Pipeline for Production AI Agents",
        "description": "An operator playbook for screening live AI agent output with an LLM judge: a confidence/stakes routing architecture to a human queue, continuous live-traffic rubric design, and per-bias mitigations for position, verbosity, and self-preference.",
        "kind": "sub",
        "order": 2,
        "html": "https://changegamer.ai/articles/llm-as-judge-screening-in-production",
        "markdown": "https://changegamer.ai/articles/llm-as-judge-screening-in-production.md",
        "json": "https://changegamer.ai/api/articles/llm-as-judge-screening-in-production.json"
      },
      {
        "slug": "agent-cost-telemetry-in-production",
        "title": "How to Track AI Agent Costs in Production",
        "description": "How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.",
        "kind": "sub",
        "order": 3,
        "html": "https://changegamer.ai/articles/agent-cost-telemetry-in-production",
        "markdown": "https://changegamer.ai/articles/agent-cost-telemetry-in-production.md",
        "json": "https://changegamer.ai/api/articles/agent-cost-telemetry-in-production.json"
      },
      {
        "slug": "incident-to-eval-fixture-loop",
        "title": "How to Turn an AI Agent Incident into an Evaluation Test Case",
        "description": "How to turn an AI agent incident into an evaluation test case: freeze the trace, redact it, minimize it to the failing decision, label the expected outcome with a trusted oracle, prove it fails then passes, and retire it later.",
        "kind": "sub",
        "order": 4,
        "html": "https://changegamer.ai/articles/incident-to-eval-fixture-loop",
        "markdown": "https://changegamer.ai/articles/incident-to-eval-fixture-loop.md",
        "json": "https://changegamer.ai/api/articles/incident-to-eval-fixture-loop.json"
      },
      {
        "slug": "retrieval-attribution-logging-in-production",
        "title": "How to Log RAG Retrieval in Production for Debugging Agent Answers",
        "description": "How to log RAG retrieval in production so a wrong agent answer can be debugged: per-retrieval fields, what reached the context window versus what the model cited, privacy by hashes and IDs, and retention and sampling choices.",
        "kind": "sub",
        "order": 5,
        "html": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production",
        "markdown": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production.md",
        "json": "https://changegamer.ai/api/articles/retrieval-attribution-logging-in-production.json"
      },
      {
        "slug": "online-feedback-signals-for-ai-agents",
        "title": "How to Measure AI Agent Quality from Live User Feedback",
        "description": "How to measure AI agent quality from live user feedback: why explicit ratings are sparse and biased, how re-asks, abandonment and escalations mislead, and how to join each signal to a trace and route it to review.",
        "kind": "sub",
        "order": 6,
        "html": "https://changegamer.ai/articles/online-feedback-signals-for-ai-agents",
        "markdown": "https://changegamer.ai/articles/online-feedback-signals-for-ai-agents.md",
        "json": "https://changegamer.ai/api/articles/online-feedback-signals-for-ai-agents.json"
      },
      {
        "slug": "online-agent-metrics-and-drift-monitoring",
        "title": "How to Detect Quality Drift in a Production AI Agent",
        "description": "How to detect quality drift in a production AI agent: baseline aggregate signals, alert on a diff against the baseline, and separate a prompt, model or tool-version change from a shift in traffic mix.",
        "kind": "sub",
        "order": 7,
        "html": "https://changegamer.ai/articles/online-agent-metrics-and-drift-monitoring",
        "markdown": "https://changegamer.ai/articles/online-agent-metrics-and-drift-monitoring.md",
        "json": "https://changegamer.ai/api/articles/online-agent-metrics-and-drift-monitoring.json"
      },
      {
        "slug": "sampling-and-retention-of-agent-traces",
        "title": "How to Sample and Retain Production AI Agent Traces",
        "description": "How to sample and retain production AI agent traces: head vs tail sampling, keep-all-errors plus a random baseline, whole-trace decisions for multi-agent runs, retention tiers and redaction before the clock starts.",
        "kind": "sub",
        "order": 8,
        "html": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces",
        "markdown": "https://changegamer.ai/articles/sampling-and-retention-of-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/sampling-and-retention-of-agent-traces.json"
      },
      {
        "slug": "crash-safe-run-start-records-for-agent-traces",
        "title": "How to Keep Trace Data When an AI Agent Crashes Mid-Run",
        "description": "How to keep trace data when an AI agent crashes mid-run: write a small run-start record outside the trace buffer, detect orphans, count them as unknown outcomes, and force-decide on shutdown.",
        "kind": "sub",
        "order": 9,
        "html": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces",
        "markdown": "https://changegamer.ai/articles/crash-safe-run-start-records-for-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/crash-safe-run-start-records-for-agent-traces.json"
      },
      {
        "slug": "choosing-an-llm-observability-backend",
        "title": "How to Choose an LLM Observability Platform for AI Agents",
        "description": "How to choose an LLM observability platform: decide on OTel-native ingestion, self-host versus cloud, export portability, redaction hooks and retention support before comparing vendors.",
        "kind": "sub",
        "order": 10,
        "html": "https://changegamer.ai/articles/choosing-an-llm-observability-backend",
        "markdown": "https://changegamer.ai/articles/choosing-an-llm-observability-backend.md",
        "json": "https://changegamer.ai/api/articles/choosing-an-llm-observability-backend.json"
      },
      {
        "slug": "redacting-sensitive-data-from-agent-traces",
        "title": "How to Redact PII From AI Agent Traces: Placement, Testing and Cleanup",
        "description": "How to redact PII from AI agent traces in practice: where the redactor sits in the pipeline, what to do per span field, how to test it with seeded fake PII, and how to clean up after a leak.",
        "kind": "sub",
        "order": 11,
        "html": "https://changegamer.ai/articles/redacting-sensitive-data-from-agent-traces",
        "markdown": "https://changegamer.ai/articles/redacting-sensitive-data-from-agent-traces.md",
        "json": "https://changegamer.ai/api/articles/redacting-sensitive-data-from-agent-traces.json"
      }
    ]
  },
  "navigation": {
    "pillar": {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "kind": "pillar",
      "order": 0,
      "html": "https://changegamer.ai/articles/agent-observability-and-evaluation",
      "markdown": "https://changegamer.ai/articles/agent-observability-and-evaluation.md",
      "json": "https://changegamer.ai/api/articles/agent-observability-and-evaluation.json"
    },
    "previous": {
      "slug": "agent-cost-telemetry-in-production",
      "title": "How to Track AI Agent Costs in Production",
      "description": "How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.",
      "kind": "sub",
      "order": 3,
      "html": "https://changegamer.ai/articles/agent-cost-telemetry-in-production",
      "markdown": "https://changegamer.ai/articles/agent-cost-telemetry-in-production.md",
      "json": "https://changegamer.ai/api/articles/agent-cost-telemetry-in-production.json"
    },
    "next": {
      "slug": "retrieval-attribution-logging-in-production",
      "title": "How to Log RAG Retrieval in Production for Debugging Agent Answers",
      "description": "How to log RAG retrieval in production so a wrong agent answer can be debugged: per-retrieval fields, what reached the context window versus what the model cited, privacy by hashes and IDs, and retention and sampling choices.",
      "kind": "sub",
      "order": 5,
      "html": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production",
      "markdown": "https://changegamer.ai/articles/retrieval-attribution-logging-in-production.md",
      "json": "https://changegamer.ai/api/articles/retrieval-attribution-logging-in-production.json"
    }
  },
  "resources": [
    {
      "slug": "testing-ai-agents",
      "html": "https://changegamer.ai/resources/testing-ai-agents",
      "markdown": "https://changegamer.ai/resources/testing-ai-agents.md",
      "json": "https://changegamer.ai/api/resources/testing-ai-agents.json"
    },
    {
      "slug": "evaluating-ai-agents",
      "html": "https://changegamer.ai/resources/evaluating-ai-agents",
      "markdown": "https://changegamer.ai/resources/evaluating-ai-agents.md",
      "json": "https://changegamer.ai/api/resources/evaluating-ai-agents.json"
    },
    {
      "slug": "agent-observability",
      "html": "https://changegamer.ai/resources/agent-observability",
      "markdown": "https://changegamer.ai/resources/agent-observability.md",
      "json": "https://changegamer.ai/api/resources/agent-observability.json"
    }
  ]
}