Guides
Long-form guides on making a website work for AI agents: visibility, access control, formats and monetization. Each guide is a pillar article plus in-depth sub-articles.
The agent-ready web
How to make a website readable, citable, controllable and payable for AI agents and AI crawlers — the operator side of the machine-first web.
- llms.txt and the Agent-Ready Website Playbook How to make a site work for AI agents: llms.txt, crawler access, structured data, licensing and machine-payable access, with a 30-day plan.
- llms.txt vs robots.txt vs sitemap.xml: Which File Does What The three root-level files every agent-ready site publishes, what each one is actually for, and why publishing one does not substitute for the others.
- How to Write an llms.txt File (Format, Template, and Maintenance) A step-by-step guide to writing a useful llms.txt: the exact format, a copy-paste template, what to put under ## Optional, how to validate it, and how to keep it from rotting.
- Serving Markdown Variants to AI Agents: The Cheapest Win in AI Visibility How to publish a .md twin of every page — URL patterns, content negotiation, discovery headers, generation pitfalls — and why it cuts what an agent pays to read you.
- How AI Search Engines Choose Sources (And What You Can Actually Influence) What is known, what is claimed and what is speculation about how ChatGPT, Perplexity and AI Overviews pick the pages they cite — and the short list of things a site owner can actually control.
- Should You Block AI Crawlers? A Decision Framework by Business Model Blocking AI crawlers is four separate decisions, not one. A framework that maps each crawler class to what it costs and earns you, by business model, with the exact robots.txt for each answer.
- What to Charge AI Crawlers: Pricing Models for Machine Buyers Per-crawl, per-resource, corpus licence or subscription key — the four ways to price AI access, the arithmetic behind each, and why pricing before you have demand data is the standard mistake.
- Implementing an HTTP 402 Paywall an Agent Can Actually Pay A working implementation guide for machine-payable content: the 402 response body, Link headers, key issuance and validation, caching rules, and the mistakes that make a 402 gate unpayable.
- Structured Data for AI Agents: Which Schema.org Types Earn Their Keep Most schema.org markup is invisible to machine readers. The types that are worth the effort for AI agents, how to emit them without drift, and what to build instead of more markup.
- JSON API Design for AI Agents: Endpoints They Prefer Over Scraping How to publish read-only JSON endpoints that agents choose over scraping your HTML: discovery index, stable shapes, freshness signals, bulk exports, and errors a machine can act on.
- Measuring AI Agent Traffic: Server-Side Telemetry That Answers Real Questions Why client-side analytics miss AI agents entirely, the minimum row schema to log, the five queries worth running, and how to tell a real crawler from a spoofed user agent.
- Licensing Content for AI Training: RSL, Terms, and Provenance How to publish machine-readable licence terms for AI use — what RSL is, what it does and does not do, how it differs from robots.txt and Content Signals, and where provenance standards fit.
- Running an MCP Server as a Distribution Channel for Your Content Why a content site should expose an MCP server, which tools to ship, how discovery and authentication work, how to gate paid tools, and the honest limits of the channel.
- Why AI Agents Can't Read Your Site: Twelve Failure Modes and How to Find Them A diagnostic catalogue of the twelve reasons AI agents and crawlers fail on real sites — from silent WAF blocks to JS-only rendering — each with the command that detects it and the fix.
Selling to AI agents
How to sell content, APIs and tools to buyers that are software — discovery, machine-readable offers, spend ceilings, payment rails, and what breaks.
- x402 Protocol and the Guide to Selling to AI Agents The operator playbook for the x402 protocol and selling content, APIs and tools to AI agents that discover, evaluate and pay on their own.
- Agent Checkout vs. Human Checkout: Why Your Payment Flow Fails Machine Buyers Why checkout built for a person watching a screen is unusable by an AI agent, and what a checkout flow that actually completes for a machine buyer looks like — 402 + API key versus native x402.
- Machine-Readable Pricing Pages: How to Let an Agent Evaluate Your Offer Before It Pays Why a prose pricing page cannot be evaluated by an AI agent, what fields a machine-readable offer catalog needs, and how to keep it in lockstep with your human pricing page and your 402 body.
- ACP vs. AP2 vs. x402: Which Agent Payment Rail Should You Implement? A decision framework for choosing between ACP, AP2, and x402 (plus the self-hosted 402 gate) — sorted by who your buyer actually is, what you are selling, and what is live versus waitlisted as of July 2026.
- How to Accept x402 Stablecoin Payments: A Seller Implementation Guide A build guide for sellers who have already decided x402 is the right rail: the 402 response shape, the wallet/facilitator/network choices, the verify-then-settle retry flow, exact vs. upto pricing, and how to ship it dormant until you are ready to go live.
- Issuing API Keys to AI Agents Automatically: A Build Guide How to design a system that mints and delivers API keys to agent and software buyers with minimal human friction: trigger models, storage, delivery, key format, tiering, rotation and revocation — illustrated with ChangeGamer's own Stripe-webhook mechanism.
- Pricing Tiers for API and Corpus Access: What Actually Varies Between Them The axes that actually distinguish one pricing tier from another for a machine buyer — rate limits, content scope, deliverables and license grant — and how ChangeGamer structures its own four tiers around deliverable and license, not gated content.
- Agent Spend Limits and Trust: What a Seller Should Verify Before Granting Access The seller-side counterpart to agent spend controls — how an API operator reads an inbound agent's spend ceiling before granting access, which payment protocols actually prove that ceiling, how to revoke access, and what "trust" operationally means for a seller when no portable agent-reputation standard exists yet.
- Refunds and Disputes with Agent Buyers: What a Seller Actually Does What happens on the seller side when an autonomous agent's purchase needs to be reversed or is disputed — API-key refund mechanics, why x402 settlement cannot be undone, what card-token revocation does and does not prove, and what to log before you reverse anything.
- Packaging a Corpus as a Product: Format, Schema, Versioning and Delivery The packaging decisions behind selling a content corpus as a dataset product — export format, the free-sample/gated-full split, a per-record metadata schema, a corpus version number, and which of three delivery mechanisms to use — grounded in ChangeGamer's own three real export formats.
- How Do AI Agents Discover Paid APIs? A Guide to Every Surface How an AI agent finds out a paid API or resource exists before it ever reads a price: llms.txt, the JSON API index, MCP registries, incidental 402 discovery, x402 auto-listing, and what .well-known does and does not cover.
- Fraud and Abuse from AI Agent Traffic: What a Seller Should Detect How a seller of APIs, content, or tools to AI agents spots and mitigates abuse once access is already granted — key sharing, over-scope scraping, spend-ceiling circumvention, spoofed identity, and rate-limit evasion patterns specific to autonomous agents.
- Measuring Revenue from AI Agent Traffic: Beyond the Traffic Log The revenue-layer fields and queries a seller adds on top of a general traffic log — authorized-vs-settled, revenue per rail, revenue per tier, and how to avoid double-counting a webhook retry as two sales.
MCP in practice
How to build, ship and run an MCP server in production — transport, auth, tool design, versioning, testing, distribution, observability, cost and failure modes.
- MCP Authentication and the Production MCP Server Playbook MCP authentication with OAuth 2.1, plus transport, tool design, versioning, testing, distribution and observability for a production MCP server.
- stdio vs. Streamable HTTP for MCP Servers: A Decision Framework Which MCP transport to build against and why: the single-client-vs-shared decision rule, how state works without a session handshake under the 2026-07-28 spec, the auth-model switching cost, and what actually breaks migrating off HTTP+SSE.
- How to Implement OAuth 2.1 for an MCP Server A wire-level implementation walkthrough for OAuth 2.1 on a remote MCP server: what the discovery documents actually contain, CIMD vs. Dynamic Client Registration in your server code, per-SEP detail from the 2026-07-28 hardening set, and token-validation mechanics.
- Defending MCP Clients Against Tool Description and Output Injection Two distinct MCP injection surfaces — a tool description at connect-time and a tool's return value at call-time — and the client-side architectural patterns (Dual LLM, Action-Selector, Context-Minimization) that contain each one.
- How to Test an MCP Server in CI The implementation mechanics below the three-layer test pyramid: what a mocked MCP transport actually replaces, what a Streamable HTTP cassette contains, a concrete CI job/trigger shape, and how to catch spec-version drift before it reaches production.
- MCP Server Versioning and Spec Migration: An Operator Playbook A migration runbook for MCP server operators: feature-detecting via capabilities instead of hard protocolVersion branching, a dual-version fleet rollout with rollback triggers, a compatibility shim for legacy clients still sending initialize, and a deprecation calendar built off the 12-month SEP-2577 floor.
- MCP Server Observability with OpenTelemetry: Spans, Metrics, and Trace Correlation Instrumenting an MCP server past the pillar's baseline: what to put on a tool-call span beyond gen_ai.tool.name, what replaces the deprecated Logging primitive in practice, per-tool-name latency and error-rate metrics, and how a trace ID actually survives the agent-to-upstream-API hop.
- How to Publish an MCP Server to the Official Registry A step-by-step walkthrough of the mcp-publisher CLI and the server.json manifest for publishing an MCP server to registry.modelcontextprotocol.io, how to republish after a version bump, and how the registry relates to aggregators, marketplaces, and direct distribution.
- MCP Server Cost Optimization: Toolset Size, Caching Hints, and Fan-Out How the token cost of an MCP server's tool list, the 2026-07-28 spec's ttlMs/cacheScope caching hints, fan-out from callers you do not control, and per-tool-name cost visibility each shape what a production MCP server actually costs to run.
- Common MCP Server Failure Modes and How to Fix Them A runtime playbook for the two MCP server failure modes with no dedicated deep-dive elsewhere: unrecoverable state after a mid-call crash, and malformed or hallucinated tool calls that reach the handler despite upstream validation.
- MCP Tools vs Resources vs Prompts: How to Choose the Right Primitive A decision procedure for MCP's three server-side primitives — who controls each one, a worked example of what it costs to expose a Resource as a Tool by mistake, and how Sampling and Elicitation fit as the client-side counterparts.
- The MCP Server Production Launch Checklist A phase-by-phase go/no-go checklist for launching an MCP server: checkable gate conditions for transport and auth, tool design, cross-client testing, publish readiness, observability, and ongoing operation — with links to the mechanics each gate depends on.
- Zero-Touch Enterprise Authorization for MCP Servers: ID-JAG and SEP-990 How Enterprise-Managed Authorization (SEP-990) removes the per-server OAuth consent screen for MCP servers: the ID-JAG grant mechanism, its RFC 8693/7523 building blocks, named launch adopters as of August 2026, and how it layers on top of standard OAuth 2.1 rather than replacing it.
RAG in production
How to run retrieval-augmented generation as a real system — ingestion and chunking, embedding choice and reindexing, hybrid search, reranking, evaluation, freshness, access control, cost, latency, and when RAG is the wrong answer.
- Agentic RAG in Production: The Complete Operator Guide The operator playbook for agentic RAG in production: ingestion, chunking, hybrid retrieval, reranking, evaluation, freshness and cost.
- How to Build a RAG Ingestion Pipeline That Survives Production The six properties that separate a production RAG ingestion pipeline from a demo script: tested extraction, idempotent writes, incremental updates, deletion propagation, durable execution, and metadata captured at ingest time.
- How to Chunk Documents for RAG (Strategy Beats Size) Chunking decisions that actually move retrieval quality: structural boundaries before fixed windows, parent-document expansion, special handling for tables and code, overlap trade-offs, and tuning against recall@k instead of blog defaults.
- How to Choose an Embedding Model for RAG (and Version It Like a Schema) Embedding selection as an operations problem: the criteria that dominate total cost of ownership, the never-mix-spaces invariant, reindex migrations with dual indexes, and where quantization fits once cost shows up.
- How to Combine Keyword and Vector Search in RAG Why hybrid retrieval is the production default rather than an upgrade: complementary failure modes of lexical and dense search, reciprocal rank fusion versus weighted scoring, parameter choices, and how filtering interacts with fusion.
- When and How to Rerank Retrieved Documents in RAG Reranking as a budget decision: why first-stage ranking misorders good evidence, when cross-encoder reranking pays for itself, how to pick candidate depth at the knee, gating by query difficulty, and deduplicating after fusion.
- How to Evaluate a RAG System (Retrieval Metrics, Generation Metrics, CI Gates) The evaluation harness that keeps RAG changeable: golden-set construction, retrieval metrics separated from generation metrics, LLM-as-judge screening with human acceptance, CI regression gates, and the logged-query flywheel.
- How to Keep a RAG Index Fresh (Staleness Bounds, Not Vibes) Freshness as an engineered property: per-source staleness contracts, document versioning and tombstones, effective-date filtering, deletion propagation with reconciliation backstops, and the sync-lag metrics that predict stale answers before users report them.
- Permission-Aware Retrieval: Multi-Tenant RAG Without Leaks How to enforce access control inside retrieval for multi-tenant RAG: authorization as mandatory pre-filters derived from the authenticated principal, tenant isolation mechanics, permission-negative testing, and the post-filter trap that produces empty answers.
- Retrieval as a Tool: Agentic RAG Patterns That Survive Production When AI agents consume retrieval as a tool rather than a pipeline stage: narrow tool contracts, hard budgets on iterative retrieval, query decomposition, provenance-carrying results, and keeping trust boundaries inside the retrieval path.
- How to Cut RAG Cost and Latency Without Cutting Quality RAG cost and latency as engineered budgets: where the money actually goes, caching layers and their hit-rate economics, routing queries to right-sized models, bounding retrieval fan-out, the hidden lines (reindex migrations, eval compute), and p95 discipline.
- Common RAG Failure Modes and How to Fix Them An operator runbook for the four RAG failure classes with no dedicated deep-dive elsewhere: retrieval miss, context overload, injection via content at ingestion time, and silent quality degradation — symptom, first diagnostic, and fix for each.
- GraphRAG vs Vector RAG: When to Use a Knowledge Graph Instead A decision framework for choosing graph-structured retrieval over standard vector RAG: which query types GraphRAG actually wins, what building a knowledge graph costs, named implementations, and hybrid vector-plus-graph patterns.
- When Is RAG the Wrong Answer? A Decision Guide A worked decision guide for the four real alternatives to retrieval-augmented generation: including knowledge directly, querying structured data with text-to-SQL, fine-tuning for behavior change, and graph-based retrieval for entity relationships.
Agent reliability in production
How to make an AI agent reliable — tool-calling contracts, structured outputs, retries and idempotency, timeouts, durable execution, guardrails, evaluation in CI, observability, incident response, and rollout.
- Agent Guardrails and the AI Agent Reliability Playbook Agent guardrails plus the eleven other disciplines that make an AI agent reliable in production: tool calling, retries, durable execution and rollout.
- How to Make AI Agent Tool Calling Reliable An operator playbook for tool-calling contracts: per-provider strict-mode config (OpenAI, Anthropic, Gemini, self-hosted grammars), the full failure-mode-to-fix table, what BFCL actually measures, and how to stop parallel tool calls from breaking a dependency chain.
- Structured Outputs vs Tool Calling: When to Use Each A decision framework for structured outputs versus tool calling in AI agents, with runnable JSON Schema examples for OpenAI, Anthropic, Gemini, vLLM/SGLang, and llama.cpp, plus mitigation code for truncation, refusal, and grammar-compilation latency.
- How to Make AI Agent Retries Idempotent A deep-dive on retrying agent tool calls safely: the transient-vs-terminal decision, why an agent side effect can fire before a failure signal reaches the caller, idempotency-key mechanics (run ID + step index), the unknown-outcome edge case, and where idempotency keys do not reach.
- When Do AI Agents Need Durable Execution? A deep-dive on durable execution for AI agents: the persisted event log, the replay-determinism constraint, the four architectural shapes mapped across ten engines and frameworks, and a decision framework for when a durable execution engine is worth adding at all.
- How to Design Guardrails for AI Agent Reliability An operator playbook for reliability guardrails: the three checkpoints (input, output, action), layering cheap checks under slow ones with a fail-closed default, the two-of-three-properties rule for when a tool call needs human approval, and logging every verdict against the run trace ID.
- How to Evaluate AI Agents in CI An operator playbook for gating an AI agent release in CI: why agent eval needs trajectory-level scoring across the tasks it actually runs, how public benchmarks diverge as proxies, ground-truth vs LLM-as-judge tool-call scoring, and the three-layer test pyramid that keeps CI fast and non-flaky.
- How to Roll Out a New AI Agent Version Safely An operator playbook for shipping a new agent version without breaking production: in-repo vs. registry prompt storage, a version-numbering comparison, the six-step promotion flow, A/B-test mechanics, the composite-version trace fields, and a rollback drill.
- How to Build an Incident Response Runbook for AI Agent Failures An operator playbook for the moment an AI agent fails in production: a triage step to classify the failure fast, trace freezing before rollback destroys the evidence, first-response depth on the four failure classes, and a blameless post-mortem that feeds back into guardrails and eval.
- How to Set Timeouts for AI Agent Tool Calls A deep-dive on timeout and deadline design for AI agents: sizing LLM-call, tool-call, and sub-agent-hop timeouts differently, allocating a wall-clock budget across a multi-step chain, and propagating a remaining-deadline value from parent to child calls.
- How to Design a Circuit Breaker for AI Agents A deep-dive on the circuit breaker pattern for AI agents: the Closed/Open/Half-Open state machine with a worked open-source example, where to place a breaker in an agent's call path, and degraded-mode fallback design as its own discipline per dependency type.
- What Should an AI Agent's Observability System Capture? An operator playbook for instrumenting an AI agent: the trace/span model behind a run, the OpenTelemetry GenAI attributes that name each field, the fields worth capturing per span, and what to redact before any of it gets logged.
- The AI Agent Production Reliability Checklist A go/no-go checklist that turns the agent reliability pillar's twelve-discipline closing list into checkable gates — the specific artifact, header, or trace field that proves each one holds, with a link to whichever sibling article owns its mechanics.
Agent security operations
How to defend an AI agent deployment from the operator side — secrets and credential hygiene, prompt-injection defense in depth, sandboxing, supply-chain provenance, least privilege, audit trails, incident response, and rate/abuse controls — not buyer-side fraud and not reliability-framed guardrails.
- How to Secure AI Agents in Production Credential hygiene, prompt-injection defense in depth, sandboxing choices for code execution, supply-chain provenance, least privilege, audit trails, incident response, and rate/abuse controls — eight operator-side defenses against an adversarial actor or a compromised dependency, not against ordinary load or failure.
- How to Manage Secrets for AI Agents in Production Why single-agent credential issuance is not the whole secrets problem: auditing which tool servers and MCP connectors hold credentials on an agent's behalf, treating provider-side prompt caches as a disclosure surface, and the named frameworks — OWASP's Secrets Management Cheat Sheet, Twelve-Factor config, and the OWASP GenAI project — that govern the rest.
- How to Defend an AI Agent Against Prompt Injection A decision framework for matching Action-Selector, Plan-Then-Execute, Dual LLM, and the other named architectural patterns to a task's actual blast radius, including when to compose two patterns together and when none of them is worth the overhead.
- How to Choose a Sandbox for AI Agent Code Execution A two-axis framework — trust in the code's source crossed with the blast radius of a successful escape — for picking an isolation layer, choosing among six hosted sandbox APIs, and hardening the harness around whichever one you pick.
- How to Verify Supply-Chain Provenance for AI Agent Dependencies An operational playbook for three separate trust-boundary gates — package-install time, model-load time, and MCP-server-connect time — that turns SBOM and attestation formats into checks a pipeline can actually run, plus a fail-closed default for the dependency that carries neither.
- How to Enforce Permission Boundaries for AI Agents at Runtime How a least-privilege grant actually gets enforced once an agent is running — the eight named checkpoints, five verdicts, and fail-closed contract in Microsoft's draft Agent Control Specification (ACS), and what happens when the policy engine that enforces the boundary breaks.
- How to Build an Audit Trail for an AI Agent A field-by-field forensic playbook for an AI agent audit log: why each of the checklist's seven required fields matters for reconstruction, a worked incident walkthrough, and the real tension between OpenTelemetry's redact-by-default tracing and full audit logging.
- How to Respond When an AI Agent Takes a Harmful Action Why the pillar's cut-credential, stop-instance, export-evidence order is structurally forced rather than a tidy convention, a concrete answer for who holds standing revocation authority, a pre-incident drill for the evidence-export path, and what changes when the compromised credential belongs to a third-party tool server instead of the agent itself.
- How to Rate-Limit and Cap Spend for Your Own AI Agent Enforcement mechanics for the two ceilings an agent operator should set before production: where a per-credential tool-call counter has to live to stay correct under concurrent calls, where a spend ceiling gets checked in the tool-call loop, and how to reject out-of-scope tool-call arguments with canonicalization rather than a naive prefix match.
- How to Verify Content Provenance for AI Agents with C2PA A three-state decision procedure — valid manifest, invalid signature, absent manifest — for what an AI agent's ingestion pipeline should do differently with a web image, an email attachment, or a retrieved document, plus a checklist for wiring a C2PA reader library into that pipeline as a gate before content reaches the model.
- How AI Agents Prove Identity and Delegated Authority The two-layer model an autonomous agent needs to pass before any credential-custody or permission question even applies: a cryptographic workload identity proving what it is (SPIFFE/SPIRE, cloud workload identity federation) and a separate delegated-authority grant proving it may act on a human's or org's behalf (OAuth scopes, RFC 8693 token exchange, RFC 8707 audience binding).
- How to Protect PII and Personal Data in AI Agent Pipelines Why an AI agent expands PII exposure past a bounded API call — large ingested context, external tool calls, persistent memory and logs, provider training risk — and the containment controls, provider data-handling terms, and GDPR/EU AI Act/CCPA compliance boundary that follow from it.
- The AI Agent Security Checklist A go/no-go checklist that turns the agent security operations pillar's eight disciplines, plus content provenance, agent identity, and data privacy, into checkable gates — the specific inventory row, test result, or logged decision that proves each one holds, with a link to whichever sibling article owns its mechanics.
Agent observability and evaluation
How to observe and evaluate an AI agent already live in production — tracing spans for tool calls and retrieval steps, judge-based screening versus human acceptance of live output, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures — not the pre-ship, CI-gating evaluation methodology already owned by evaluating-ai-agents-in-ci, and not the trace/span field mechanics already owned by agent-observability-for-reliability, both in the completed agent-reliability cluster.
- AI Agent Observability and the Production Evaluation Playbook AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.
- Distributed Tracing for Multi-Agent AI Systems How trace ID propagation actually works across a multi-agent handoff, why handoff and delegation need different trace shapes, and why an orchestrator-level trace can hide a failed leg of a fan-out.
- Designing an LLM-as-Judge Pipeline for Production AI Agents An operator playbook for screening live AI agent output with an LLM judge: a confidence/stakes routing architecture to a human queue, continuous live-traffic rubric design, and per-bias mitigations for position, verbosity, and self-preference.
- How to Track AI Agent Costs in Production How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.
- How to Turn an AI Agent Incident into an Evaluation Test Case How to turn an AI agent incident into an evaluation test case: freeze the trace, redact it, minimize it to the failing decision, label the expected outcome with a trusted oracle, prove it fails then passes, and retire it later.
- How to Log RAG Retrieval in Production for Debugging Agent Answers How to log RAG retrieval in production so a wrong agent answer can be debugged: per-retrieval fields, what reached the context window versus what the model cited, privacy by hashes and IDs, and retention and sampling choices.
- How to Measure AI Agent Quality from Live User Feedback How to measure AI agent quality from live user feedback: why explicit ratings are sparse and biased, how re-asks, abandonment and escalations mislead, and how to join each signal to a trace and route it to review.
- How to Detect Quality Drift in a Production AI Agent How to detect quality drift in a production AI agent: baseline aggregate signals, alert on a diff against the baseline, and separate a prompt, model or tool-version change from a shift in traffic mix.
- How to Sample and Retain Production AI Agent Traces How to sample and retain production AI agent traces: head vs tail sampling, keep-all-errors plus a random baseline, whole-trace decisions for multi-agent runs, retention tiers and redaction before the clock starts.
- How to Keep Trace Data When an AI Agent Crashes Mid-Run How to keep trace data when an AI agent crashes mid-run: write a small run-start record outside the trace buffer, detect orphans, count them as unknown outcomes, and force-decide on shutdown.
- How to Choose an LLM Observability Platform for AI Agents How to choose an LLM observability platform: decide on OTel-native ingestion, self-host versus cloud, export portability, redaction hooks and retention support before comparing vendors.
- How to Redact PII From AI Agent Traces: Placement, Testing and Cleanup How to redact PII from AI agent traces in practice: where the redactor sits in the pipeline, what to do per span field, how to test it with seeded fake PII, and how to clean up after a leak.