ChangeGamer

← All guides · Agent observability and evaluation

How to Log RAG Retrieval in Production for Debugging Agent Answers

Part 5 of Agent observability and evaluation · 1,543 words · ~7 min read · published 2026-10-01 · updated 2026-10-01 · Markdown variant

How to log RAG retrieval in production so a wrong agent answer can be debugged: per-retrieval fields, what reached the context window versus what the model cited, privacy by hashes and IDs, and retention and sampling choices.

In short

  • A production RAG retrieval log should record, per retrieval span, the query as issued and as rewritten, the index and its version, the returned document IDs with scores, and chunk-text hashes.
  • Logging only what a retrieval step returned is not enough for debugging, because an operator also needs to know which chunks actually reached the model context window and which the answer cited.
  • The purpose of a retrieval log is to show which runbook branch applies to a bad answer: a retrieval miss, where nothing relevant came back, or ignored context, where relevant text was present but unused.
  • Retrieval logs should store document IDs and content hashes by default instead of raw chunk text, so the log stays useful for debugging without becoming a second copy of sensitive source data.
  • No resource in the ChangeGamer corpus names retrieval-attribution logging as a practice, so the field list in this article is reasoned design built on the span model, not a published standard.

Part of the AI Agent Observability and the Production Evaluation Playbook guide.


A retrieval log earns its place by letting you say, for one bad answer, which part of the pipeline failed. The pillar defines retrieval-attribution logging in one section; this article designs the log itself, field by field. It stops at the point where diagnosis hands off: causes and fixes belong to the RAG failure modes runbook, and metrics and golden sets belong to evaluating RAG systems. Span attribute names are covered in agent observability for reliability.

One hedge applies throughout. No resource in the ChangeGamer corpus names retrieval-attribution logging. The agent-observability resource says a trace is one run and a span is one step, including a retrieval, and the fields below are reasoned design on top of that model as of 1 October 2026, not a standard.

What should one retrieval span record?

One retrieval span should record the query, the index, the candidates returned, and the fate of each candidate, because those four answers are what a post-incident reader needs. The fields below are a starting set, not a schema anyone has published.

Field Why it earns a place
trace_id, parent span ID Links the retrieval to the run and to the model call that consumed it
Query as issued What the caller or agent actually asked
Query as rewritten What was sent to the index after rewriting or decomposition
Index name and index version Tells you which snapshot answered, and whether it was the one you meant
Embedding model identifier Query and document vectors must come from the same model
Returned document IDs, rank, score The candidate list, in order, with the number that ranked it
Chunk-text hash per candidate Proves which exact text version was returned without storing the text
Reranker output order, if any Shows whether reranking promoted or buried the relevant chunk
Placed in context (yes/no, position) Separates "retrieved" from "the model saw it"
Cited in answer (yes/no) Separates "the model saw it" from "the model used it"

The rag-retrieval-for-agents resource is the reason the rewritten-query and reranker rows exist. It describes agentic retrieval where a model rewrites or decomposes the query, and a pipeline that fetches a wide candidate set and reranks it to a small one. Each of those steps changes what the model sees, so each needs a trace of its output.

Why log both the issued and the rewritten query?

Log both because a rewrite can lose the very term the answer depends on, and only the pair shows it. The rag-retrieval-for-agents resource describes query rewriting and decomposition as a way to address vocabulary mismatch and multi-part questions. A rewrite is also a model call that can drop a qualifier or product name.

If only the rewritten query is stored, a retrieval that looks reasonable on its own can hide a faulty rewrite. If only the original is stored, you cannot tell whether the index was asked the right thing. With decomposition, record each sub-query as its own span under one parent so the fan-out stays readable.

Why record the index version and embedding model?

Record the index version and embedding model because "the index returned the wrong chunk" is unanswerable if you cannot say which index, built when, with which embedder. The embeddings-vector-search resource states that query and document vectors only work together if they come from the same model, so swapping models forces a full re-index.

A version identifier is therefore the difference between "retrieval regressed" and "retrieval ran against an index built before the source changed". If staleness is your suspect, pair this field with the checks in RAG index freshness.

What proves a chunk reached the context window?

A logged flag set at prompt-assembly time proves it, because retrieval output and prompt content are separate stages that can diverge. The rag-retrieval-for-agents resource lists assemble-context as its own pipeline stage with its own failure modes: a token-budget overflow, lost-in-the-middle placement, and contradictory chunks.

So the assembly step should write back to the retrieval span, per candidate:

Without this write-back, a chunk that was retrieved but trimmed for budget looks identical in the log to one the model read and ignored. Those two lead to different branches of the runbook.

How do you tell which chunks the model used?

Record citations when your answer format has them, and treat the result as evidence about use rather than proof. If the model is asked to cite chunk IDs, log the IDs it returned and compare them with the placed set. A cited chunk was at least referenced. An uncited chunk may still have shaped the answer.

Where the format has no citations, do not invent an attribution score and present it as fact. You can still log the placed set, which supports the question that matters most for triage: was the relevant text in front of the model at all?

What must the log show to pick the right runbook branch?

The log must show whether a chunk containing the correct answer was retrieved, placed in the context, and cited, and the first stage where that chain breaks tells you the branch. This is the entire diagnostic job of the log; what to do on each branch lives in the runbook.

What the log shows Branch
No candidate contains the answer, whatever the rank Retrieval miss: the problem is upstream of generation
A relevant candidate was returned but dropped before the prompt Assembly problem: retrieved, never seen
A relevant chunk was in the prompt and the answer contradicts or ignores it Ignored context: the problem is in generation
The relevant chunk was in the prompt and cited, and the answer is still wrong The source text or its freshness, not the pipeline

To use the table you need a way to decide that a chunk "contains the answer". For one incident, a person reading the stored text or the source document by ID is enough. At scale that judgment is evaluation work, covered in evaluating RAG systems, and the evaluating-ai-agents resource is explicit that step-level scoring pinpoints failures such as bad tool selection versus bad arguments. Retrieval deserves the same step-level view.

What should the log keep out?

By default the log should keep out raw chunk text and raw user queries, storing IDs and hashes instead. The agent-observability resource says to redact PII before logging tool-call inputs and notes that prompt and completion bodies are off by default in OpenTelemetry for PII safety. Retrieved chunks are the same kind of payload.

IDs and hashes give you most of the debugging value. The ID lets you fetch the current source text, and the hash tells you whether that text still matches what the model saw. A mismatch is itself a finding: the source changed after the run. Where you must keep text for a subset of runs, treat that store as sensitive, with its own access control. The handling rules are in data privacy and PII for AI agents.

How long should you keep retrieval logs, and should you sample?

Keep every retrieval span at the tracing layer and decide retention by what your incident process needs, since no corpus resource sets a retention window or a sampling rate. The pillar's split applies: trace everything, then sample at the screening stage, because an unlogged run cannot be investigated, scored or reviewed later.

Two practical rules follow from reasoning, not from a published source:

What does ChangeGamer itself log?

ChangeGamer logs page fetches, not retrieval attribution, and it has no retrieval step to attribute. As of 1 October 2026, the corpus export in changegamer/src/pages/api/corpus.jsonl.ts (lines 16-64) maps the static resources array to NDJSON on each request, with no query, index or ranking. The only access log is logAccess in changegamer/worker/index.ts (lines 185-196), which writes four strings to Analytics Engine: category, slug, outcome and a user agent truncated to 100 characters.

scripts/crawler-stats.mjs aggregates that log into per-crawler counts and a top-slug table, and its own output header states that a crawl "does not mean the page was used in an answer" (lines 143-146 of the script's report template). That is the same gap this article is about, seen from the publisher's side: a fetch record shows what was read, not what a model did with it. The field design above is therefore not ChangeGamer's own practice.

Frequently asked questions

What should I log for each RAG retrieval call?
Log the query as issued and as rewritten, the index name and version, the returned document IDs with their scores and ranks, a hash of each chunk text, and which chunks were placed in the prompt. Add the trace ID so the retrieval span links to its parent run.
Should I store the full retrieved text in my logs?
Not by default. Store document IDs and content hashes, and fetch the text from the source system when you need it. Raw chunk text in logs duplicates source data into a store with different access rules, so keep full text only for a deliberately chosen, access-controlled subset.
How do I tell a retrieval miss from the model ignoring good context?
Check whether a chunk containing the answer appears in the logged context for that run. If no such chunk was retrieved or placed in the prompt, it is a retrieval miss. If it was in the prompt and the answer still contradicts it, the model ignored its context.
Do I need to log every retrieval, or can I sample?
Write the retrieval span for every run at the tracing layer, because an unlogged run cannot be investigated later. Sampling belongs downstream, where you choose which logged runs get scored by a judge or reviewed by a person, as the cluster pillar describes.
Is retrieval-attribution logging an OpenTelemetry standard?
No published standard defines it. Per the agent-observability resource's July 2026 check, the OpenTelemetry GenAI conventions, which carry Development status, treat a retrieval step as a span, and the specific fields to keep on that span are a design choice you make for your own pipeline.

#agents #observability #rag #retrieval #logging #debugging #production

Put this corpus inside your own agents

This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.

Agents: this page as Markdown · JSON · offers at /api/pricing.json · payment methods at /api/payment.json · single-resource access via HTTP 402 (how that works)