How to Track AI Agent Costs in Production
How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.
- Tracking AI agent cost in production means storing raw token counts on every span and computing dollars later against a versioned price table, so a provider price change never rewrites the cost of past runs.
- Cached and uncached input tokens should be recorded as separate counts on each model-call span, because a single blended input-token number cannot reproduce the invoice once cache hits are billed at a different rate.
- Retried, failed and abandoned agent runs spend real money, so cost telemetry should attribute their spend to the task that caused it instead of dropping them from the denominator of cost per successful task.
- Trace sampling must never apply to cost accounting: a sampled span can be discarded for storage reasons, but its token counts still need to reach a cost counter, or the total will drift below the invoice.
- Budget alerts for agents work best in three layers: soft and hard thresholds on cumulative spend, a per-run kill limit that stops a runaway loop, and a rate-of-spend alert that fires before a cap is reached.
- ChangeGamer's scripts/seo-budget.mjs enforces a USD 5.00 monthly hard cap on third-party measurement API spend by pairing a pessimistic preflight authorize step with a record step that writes the provider-reported actual.
Cost tracking for a production AI agent is a bookkeeping problem before it is a monitoring problem: the question is whether the number you report can be reproduced from stored facts, and whether it will still match the invoice next month. The pillar defines cost telemetry and its aggregation axes, and agent observability for reliability covers the per-span cost field; this article covers what breaks when that field meets real billing. The mechanisms below are this article's own reasoning, not a published standard.
What should a span record so its cost can be recomputed later?
A span should record raw usage, not a finished dollar figure, so cost can be recomputed whenever the price changes. The agent-observability resource lists gen_ai.request.model, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens among the OpenTelemetry GenAI attributes, and says the GenAI vocabulary is still labeled Development rather than final; as of September 2026, treat those names as likely to shift and keep a thin mapping layer of your own.
Store three counts per model call: uncached input tokens, cached input tokens, and output tokens. Then derive dollars at query time. If you store only a computed dollar amount, every later price change leaves you unable to tell whether a cost movement came from behavior or from a rate card.
Versioning the price table so historical cost stays stable
A price table keeps historical cost stable when it is versioned with an effective date and each span is priced against the version in force when it ran. A provider that lowers a rate would otherwise make last quarter's runs look cheaper after a recompute, and a raised rate would make them look worse, so a cost trend chart would move without any change in agent behavior.
A workable layout:
- One row per model and price component, each with a start date and an end date.
- The span stores the model identifier and token counts; the cost query joins on the span's timestamp.
ChangeGamer's own scripts/seo-budget.mjs shows the failure mode on a small scale. Each price in its table carries a verified date and a source, and prices marked estimated get a safety multiplier of 2. Its comment on the aeo.llm_response entry, verified 2026-09-04, records that an earlier $0.01 figure was an 8x under-estimate, corrected after measured calls came back at roughly $0.072 and $0.02. A stale price was invisible until real actuals were compared to it.
Why record cached and uncached input tokens separately?
Cached and uncached input tokens are billed differently by the provider, so a single blended input count cannot reproduce what you will be charged. The agent cost and latency optimization resource states that every major provider applies some discount on cache hits but that the percentage and minimum prompt size differ per vendor. That vendor variance is a telemetry problem in its own right: a fixed discount assumed in one place will be wrong for at least one provider.
Recording the split also makes cache behavior observable. A falling cached share of input tokens on an unchanged prompt is a cost regression you can only see if the two counts were never merged.
Attributing retried and abandoned runs to the task that caused them
Retried and abandoned runs should be charged to the task that triggered them, and their cost must stay in the books even when the user never saw an answer. Retries, tool-call failures and runs the user abandoned all consumed tokens, so dropping them makes cost per task look better than the invoice supports.
Track three numbers instead of one:
| Metric | What it includes | What it answers |
|---|---|---|
| Cost per attempt | Every run, including retries | What did we spend? |
| Cost per completed task | All attempts' spend, divided by completed tasks | What does a delivered result cost? |
| Waste share | Spend on failed, retried or abandoned attempts | How much of the bill bought nothing? |
Define "abandoned" explicitly, because the definition changes the waste share. Attaching a retry to its parent run's trace, using the trace ID the observability resource describes, is what lets a retry's spend roll up to the original task rather than appearing as an unrelated request.
Attributing shared costs to tenants
Shared costs have no natural owner, so you must pick an allocation rule and apply it consistently. A long system prompt read by every tenant's requests, a retrieval index, or a background summarization job all serve many tenants at once. No corpus resource or published standard prescribes an allocation rule, so treat the following as options, not a recommendation:
- By request share: each tenant carries its fraction of total requests. Simple, but it under-charges a tenant whose requests are unusually large.
- By token share: allocate in proportion to tokens consumed. It tracks usage more closely, and it needs the token counts to be trustworthy.
- Unallocated overhead: keep shared spend as its own line and never distribute it. It avoids arbitrary splits, at the price of an unattributed remainder.
Tag each tenant's own model calls at the span level; shared costs are the residue. Documenting the rule matters more than which one you pick.
Keeping sampling out of cost accounting
Sampling is acceptable for stored spans but never for cost accounting, because a discarded span's tokens were still billed. Even a team that samples its stored traces should increment a cost counter from every model call before any sampling decision is made.
The practical shape is two paths from the same instrumentation point. One path emits a span that a sampler may drop. The other adds token counts to a metric counter, keyed by model and by whatever tenant or task labels you have, that no sampler touches. Keep counter labels low in cardinality and put per-run detail in the span.
How do you reconcile estimated cost against the provider invoice?
Reconcile by comparing your computed spend to the provider's billed total on a fixed cadence and treating the difference as a signal to investigate, not a rounding error. Common causes are a stale price row, an unrecorded cached-token rate, sampled-out spans, and calls made outside the instrumented path, such as a script or a notebook using the same API key.
ChangeGamer's spend governor makes this discipline explicit. Its ledger in agents/SEO-AEO-LEDGER.md states that actuals come from the provider, and that when a call returns no cost field the row is recorded as an estimate and labeled estimate in the Source column, so it is visibly weaker than its neighbors. If the ledger and the provider's billing page disagree, the provider is right and the file gets a correcting row noted as a reconciliation. For agent telemetry, mark every cost figure as measured or estimated, and let the invoice win.
Budget guards for AI agents: thresholds, kill limits and rate of spend
A budget guard has three layers: cumulative thresholds, a per-run limit, and a rate-of-spend alert. Each catches a failure the others miss.
- Soft and hard thresholds on cumulative spend per period. Crossing the soft one notifies a human; crossing the hard one stops discretionary calls.
- A per-run kill limit that aborts a single trace once its running cost passes a ceiling, which is the guard against a loop that never terminates. Fan-out makes this necessary, as the cost resource notes that sub-agent dispatch multiplies spend well beyond a single-call baseline.
- A rate-of-spend alert comparing spend over the last hour to a normal hour, which fires days before a monthly threshold would.
No corpus resource specifies threshold values, and this article does not invent them; derive yours from your own baseline.
A first-hand example: authorize, then record
ChangeGamer's scripts/seo-budget.mjs runs exactly a soft/hard threshold guard with a preflight and a ledger. To be clear about its scope, it governs third-party SEO and measurement API spend against a USD 5.00 monthly hard cap. It does not meter LLM tokens.
node scripts/seo-budget.mjs status
node scripts/seo-budget.mjs authorize --item aeo.llm_response --qty 4
# exit 0 = authorized, exit 1 = denied
node scripts/seo-budget.mjs record --sku aeo.llm_response --qty 4 --cost <provider-reported-usd> --cycle <id>
authorize adds the estimate for the requested calls to month-to-date spend and exits 1 if the projection would cross the USD 5.00 hard cap, the USD 3.80 soft cap, or a per-category sub-budget. Prices marked estimated are multiplied by 2 in that projection. record then writes the provider-reported actual, and if the ledger ends up over the cap it still writes the row but exits 1, because the ledger must stay honest. Estimates gate the call; actuals govern the books. That split transfers directly to agents: a pessimistic preflight before a costly run, measured usage after it.
Where do the cost-reduction levers live?
The levers for reducing agent spend live in the agent cost and latency optimization resource, and this article deliberately does not list them. Telemetry shows where the money went; choosing what to change is a separate discipline. For the trace-level view these numbers roll up into, return to the pillar.
Sources and further reading
Token-usage attributes and the Development status of the GenAI conventions are in agent observability and tracing. The spend-governor behavior is read directly from ChangeGamer's scripts/seo-budget.mjs and agents/SEO-AEO-LEDGER.md as of 29 September 2026.
Frequently asked questions
- Should an AI agent store cost in dollars or in token counts?
- An AI agent should store token counts, split into input, cached input and output, on each span, and compute dollars at query time from a dated price table. Stored dollar figures go stale when a provider changes its rates and cannot be recomputed, while raw token counts can always be repriced.
- Why does my AI agent cost dashboard not match the provider invoice?
- A dashboard usually drifts from the invoice because of an out-of-date price table, cached tokens billed at a different rate than the dashboard assumes, spans dropped by sampling, or spend that never passed through the instrumented code path. Reconciling a dashboard against the invoice on a fixed cadence exposes which of these it is.
- How do I attribute shared costs like a system prompt to individual tenants?
- Shared costs such as a system prompt or a retrieval index have no single owner, so teams must choose an allocation rule, for example splitting by each tenant's share of requests, and apply it consistently. No published standard prescribes one rule, so the choice is a policy decision worth documenting.
- Does OpenTelemetry cover AI agent cost tracking?
- OpenTelemetry's GenAI conventions cover the inputs to cost tracking, such as the model name and gen_ai.usage.input_tokens and gen_ai.usage.output_tokens on a span, but the price lookup and the dollar figure are computed by your own code or your observability backend. As of September 2026 the GenAI attribute vocabulary is still labeled Development rather than stable.
This guide is free and stays free. The reference corpus behind it — machine-readable contracts, verified primary sources, continuously refreshed — is the paid product: a €5 starter key unlocks every premium reference for one agent via API; a €25 corpus license delivers the full corpus as RAG / fine-tuning data with an explicit AI-use grant; the €150 enterprise license adds commercial redistribution rights.