How to Turn an AI Agent Incident into an Evaluation Test Case
How to turn an AI agent incident into an evaluation test case: freeze the trace, redact it, minimize it to the failing decision, label the expected outcome with a trusted oracle, prove it fails then passes, and retire it later.
- Turning an AI agent incident into an evaluation test case starts with freezing the production trace before any rollback, because recovery steps can destroy the only faithful record of the failing run.
- A trace must be redacted of personal data and secrets before it becomes a committed test fixture, since a fixture lives in version control far longer and is read by far more people than the original trace.
- An incident fixture should be minimized to the single failing decision, keeping only the context that changes the outcome, so the test stays readable and does not break every time an unrelated prompt detail changes.
- The expected outcome of an incident fixture should be labeled by a human or a deterministic check, not by an LLM judge alone, because a judge cannot certify a case it may itself have misjudged in production.
- A new incident fixture is only trustworthy once it has failed against the buggy version and then passed against the fix, the same red-then-green discipline used for any regression test.
- Some incidents deserve a machine-checked fixture while others are better served by a written checklist rule, and ChangeGamer's July 2026 missed-windows incident became a protocol rule plus a status-board detector rather than a build check.
Turning an incident into an evaluation test case is a sequence of small, deliberate conversions, each of which can fail quietly if skipped. The pillar explains why a postmortem should feed the suite that gates releases, and how to evaluate AI agents in CI covers the suite itself; this article covers the craft between them. The steps below are this article's own reasoning as of 30 September 2026, not a published standard.
What has to happen before you roll back?
Freeze the trace before you roll back, because the rollback can overwrite or expire exactly the evidence you need. Export the full trace tree, the model and prompt versions in force, and the tool outputs the run saw. The incident response runbooks article treats freezing as an incident step; here it doubles as fixture raw material, because a fixture built from memory of what went wrong is a guess.
The agent-observability resource is the reason this is possible at all. It describes a trace as one complete agent run under a stable trace_id, with each LLM call, tool call and retrieval as a span. A frozen trace of that shape holds the input, the decision points and the tool results in one artifact.
Why redact before the trace becomes a fixture?
Redact first because a fixture is committed to version control, where it outlives the trace store's retention and reaches far more readers. The agent-observability resource says to redact PII before logging tool-call inputs and outputs, and notes prompt and completion bodies are off by default in OpenTelemetry for PII safety. A trace exported during an incident may still carry values that a later, stricter pipeline would have removed.
Treat the frozen trace as sensitive and the fixture as the sanitized derivative:
- Replace personal data with stable placeholders so the case keeps its shape.
- Remove credentials and sensitive headers. The testing-ai-agents resource says to scrub credentials and sensitive headers from cassettes before committing, and names
filter_headersandfilter_post_data_parametersin VCR.py and pytest-recording for it. - Keep the unredacted original in restricted storage only if your retention rules allow it. No corpus resource sets a window, so that decision is yours.
How do you minimize a trace to the failing decision?
Minimize by cutting the trace down to the smallest input that still reproduces the wrong decision, then confirming the failure survives each cut. An incident trace is often dozens of steps, but the defect usually lives in one: a tool argument built wrongly, a result misread, a retry that repeated a side effect.
A workable order:
- Identify the span where the run first went wrong, not where the damage surfaced.
- Keep that span's input, the tool results that fed it, and the instructions that governed it.
- Delete everything else, re-running after each cut. If the failure disappears, restore what you removed.
Minimizing also protects the suite. A fixture stuffed with full conversation history breaks whenever an unrelated prompt line changes, which teaches the team to ignore red builds.
How should you label the expected outcome?
Label the expected outcome with a human decision or a deterministic check, not an LLM judge alone. The incident happened because something already got a case wrong in production, and if a judge was part of that screening path, reusing it as the sole oracle risks encoding the same blind spot into the test.
The evaluating-ai-agents resource lists position, verbosity and self-preference biases for LLM judges and says tool-call correctness is better measured by direct comparison against a reference answer. For an incident fixture that suggests a hierarchy:
| Oracle | Use it when | Weakness |
|---|---|---|
| Deterministic assertion on a tool call or field | The right behavior is exact, such as the correct argument or no duplicate write | Cannot judge free text |
| Human-written expected outcome | The right behavior needs judgment once, at authoring time | Costs reviewer time per case |
| LLM judge with a human-verified label | The outcome is prose and the rubric is stable | Judge can drift; label must come from a person |
The judge-design article covers building a judge; this article only says to keep a person or an exact check as the final word on a fixture.
How do you prove the fixture is red, then green?
Prove it by running the new case against the buggy version, watching it fail, then running it against the fix and watching it pass. A fixture that passes on the buggy code is asserting something the bug never violated, so it adds cost without adding protection.
# on the fix branch: take the buggy code, keep the new fixture, expect a failure
git checkout <pre-fix-commit> -- src/
pytest tests/agents -k incident_1234 # must FAIL
git checkout HEAD -- src/
pytest tests/agents -k incident_1234 # must PASS
The path and test name are placeholders for your own layout. What matters is the order: red on the old behavior, green on the new. Record both results in the pull request that adds the fixture.
Which layer should the fixture run in?
Choose the cheapest layer that can still fail on the defect: a unit test if the bug is in code around the model, a cassette replay if it depends on a recorded model exchange, and a live run only if nothing else reproduces it. The testing-ai-agents resource describes three layers: mocked-model unit tests, cassette replay on every commit, and a small live set in a nightly or pre-release job. It says most bugs live in the glue code, not the model.
That gives a decision rule for incidents:
- Argument built wrongly or output parsed wrongly: unit test with the model mocked.
- Behavior tied to one recorded model response: cassette replay. The resource says CI should be configured so a missing cassette fails instead of silently making a live call, and its cross-links note production traces can seed new cassettes and test cases.
- Behavior that only appears with the live model: a curated live scenario, kept few because the resource calls these runs costly and flaky.
Whether every trace converts cleanly to a cassette depends on your tooling; the resource states the seeding idea, not a conversion procedure.
How do you avoid an over-fitted fixture?
Avoid over-fitting by asserting the behavior that was wrong, not the exact text or path the buggy run produced. The testing-ai-agents resource advises asserting on structure and argument values rather than free-text reasoning, and asserting against a schema and key fields rather than exact prose, which tolerates benign rephrasing while catching real regressions.
Two failure modes recur. A fixture pinned to a full tool-call sequence fails whenever the agent legitimately takes a different route to the same correct result. A fixture pinned to wording fails on harmless rephrasing. Write the assertion as the invariant the incident violated, for example "the refund tool is called at most once for this order", not "the agent's third message says X".
What provenance should each fixture carry?
Each fixture should carry the incident ID, the date, the affected versions and a one-line statement of the violated invariant, so a future reader can judge whether it still matters. A case named test_case_47 becomes impossible to retire or trust.
Put the provenance next to the case, in a comment or metadata file, and link the postmortem. This is the same habit of dating claims that keeps documentation honest, applied to tests.
When should you retire an incident fixture?
Retire a fixture when the behavior it guards no longer exists, for example the tool was removed or the workflow redesigned, and record why in the same change. No corpus resource sets a retention rule for fixtures, so treat retirement as a review question, not a schedule. A suite that only grows accumulates cases that pass for reasons unrelated to their intent, and provenance is what makes the retirement decision cheap.
Which incidents deserve a fixture, and which only a rule?
An incident deserves a machine-checked fixture when its cause is a repeatable behavior a test can observe, and only a written rule when the cause is a process gap. ChangeGamer has one example of each, and both are its own history.
The 2026-07-27 and 2026-07-28 missed article windows, recorded in agents/SEO-PLAN.md under "Incident: three missed windows", produced no branch, PR or journal entry in three scheduled runs. Its root cause is unconfirmed; the leading hypothesis is that the first unit of work was too large. What came out of it was a written protocol rule, checkpoint discipline in section 3a of .claude/commands/article-cycle.md (push the branch before writing, resume an existing branch), plus a status-board check that flagged the misses. It is not a build- or CI-enforced invariant.
By contrast, src/lib/articles.ts fails the build when a sub-article does not link to its pillar. That rule is machine-checked, so a regression breaks the build instead of waiting for a person to notice.
Neither form is better in general. The test is whether a machine can observe the failure. If it can, write the fixture. If it cannot, write the rule and, where possible, a detector.
Frequently asked questions
- What is an incident fixture in AI agent evaluation?
- An incident fixture is a labeled test case built from a real production failure: the recorded input or trajectory that went wrong, paired with the correct expected outcome, added to the suite that gates future releases. It turns one incident into a permanent regression check instead of a postmortem nobody re-reads.
- Can I use a raw production trace as a test case?
- A raw production trace should not be committed as a test case without redaction and minimization. Traces can contain personal data, secrets and unrelated context, and the testing-ai-agents resource warns to scrub credentials and sensitive headers from cassettes before committing them to version control.
- How do I know an incident test case actually catches the bug?
- Run the new test against the version that caused the incident and confirm it fails, then run it against the fixed version and confirm it passes. A fixture that never failed proves nothing, because it may be asserting something the buggy version already satisfied.
- Should every AI agent incident become an automated test?
- No. An incident whose cause is a repeatable, checkable behavior is a good fixture candidate, while an incident caused by a process gap, such as a missed handoff, may be better fixed with a written rule. Choosing per incident avoids a suite full of tests that guard nothing real.
This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.