# LLM Tokenization and Token Counting: tiktoken, Provider Counting Endpoints, and Why Counts Differ

> Decision rule for counting LLM tokens: bill from the response usage fields, gate requests with the provider's token-counting endpoint, estimate offline with a tokenizer library such as tiktoken, and never reuse a count across models. Includes pitfalls and budgeting rules.

Category: Guide · Updated: 2026-10-01 · Tags: tokenization, token-counting, cost, tiktoken, context-window, agents
Canonical: https://changegamer.ai/resources/llm-tokenization-and-token-counting
Variants: [HTML](https://changegamer.ai/resources/llm-tokenization-and-token-counting) · [JSON](https://changegamer.ai/api/resources/llm-tokenization-and-token-counting.json)
License: https://changegamer.ai/license.xml · Access: free

Count tokens with the tokenizer of the exact model you will call: take billing truth from the `usage` fields of the response, gate requests before sending with the provider's token-counting endpoint when one exists, and use an offline tokenizer library (for example tiktoken for OpenAI models) only as an estimate. A token count is a property of a (text, tokenizer) pair, so a count measured for one model is not a valid number for another.

## Key facts

- **tiktoken** is described in its README as "a fast BPE tokeniser for use with OpenAI's models." `tiktoken.encoding_for_model("gpt-4o")` returns the encoding for a model; `o200k_base` and `cl100k_base` are two of the encodings it ships. The README notes BPE is reversible and lossless and works on arbitrary text, and that a token averages about 4 bytes in practice.
- **Anthropic** offers a token-counting endpoint (`POST /v1/messages/count_tokens`) that accepts the same structure as a message request, including system prompts, tools, images and PDFs, and returns `input_tokens`. Its docs call the result an **estimate** that can differ "by a small amount" from actual usage, and say you are not billed for any system-added tokens it includes.
- The endpoint is **free to use** but rate limited separately from message creation. It rejects some inputs the Messages API accepts: most server tools (web search, code execution and others), the MCP connector, and image or document blocks with a `url` or `file` source (send base64 instead).
- Anthropic documents that models from Claude Opus 4.7 onward use a newer tokenizer that produces roughly 30 percent more tokens for the same text (content-dependent) and tells you to recount against the model you plan to use.
- OpenAI's cookbook chat-counting example adds 3 tokens per message and 3 tokens to prime the reply, and warns the count "may change from model to model"; treat it as "an estimate, not a timeless guarantee."

## Decision rule

| Need | Use | Why |
|---|---|---|
| What you were actually billed or rate-limited on | The `usage` fields in the provider response | Only the provider's own accounting is authoritative |
| Gate or route a request before sending, on a hosted provider | The provider's token-counting endpoint (Anthropic: `count_tokens`) | Uses the target model's real tokenizer and request structure |
| Fast offline estimate, no network call | A tokenizer library, e.g. tiktoken for OpenAI models | Cheap and local, but an approximation of request framing |
| Model with no published tokenizer or counting endpoint | Its own tokenizer if available, plus a safety margin | Do not borrow another vendor's tokenizer as a stand-in |

Google documents a countTokens method for Gemini, but that page was not reachable when this entry was written; verify it directly before relying on it.

## Why counts differ across models

Each model family is trained with its own vocabulary and merge rules, so the same string splits into a different number of tokens. No tokenizer is universal: tiktoken is for OpenAI models, and running it on another vendor's model yields a rough proxy at best. Counts also move between versions of the same vendor, as Anthropic's tokenizer change above shows.

## How to count before sending

1. Build the real request (system prompt, tools, history, attachments) exactly as you will send it.
2. If the provider has a counting endpoint, call it with the target model ID. For Anthropic, count the same request under your current and your candidate model and compare `input_tokens` to measure a migration.
3. Otherwise, count the text with the matching offline tokenizer and add a safety margin.
4. After the real call, log the response `usage` and compare it with your estimate; tune the margin from that gap.

## Budgeting

- Input and output tokens are metered separately and priced separately by providers; budget them as two numbers, and reserve output headroom inside the context window. See /resources/context-window-management.
- Cached input is accounted differently from fresh input. Anthropic notes its counting endpoint does not apply caching logic, so it will not show cache savings; measure those from the real response. See /resources/prompt-caching-for-agents.
- Reasoning/thinking tokens can count toward input on later turns depending on the model; Anthropic documents this per model, so check before assuming history is free.
- For cost levers beyond counting, see /resources/agent-cost-latency-optimization. For picking a model by capability and context size, see /resources/choosing-an-llm-for-agents.

## Pitfalls

- **Estimate vs billed.** Offline and endpoint counts are estimates; only response usage is billing truth.
- **Multimodal tokens.** Images and PDFs consume tokens by provider-specific rules; count them through the provider rather than guessing. See /resources/multimodal-agents.
- **Tool schemas.** Tool definitions are part of the input and are counted; a text-only count that omits them under-reports. See /resources/reliable-tool-calling and /resources/structured-outputs-and-json-mode.
- **Tokenizer changes between model versions.** Recount whenever you change model ID, even within one vendor.
- **Message-framing overhead.** Chat formats add tokens around each message beyond the visible text.

## Verified sources

Fetched directly this session:

- tiktoken README (BPE tokeniser for OpenAI models, `encoding_for_model`, `o200k_base`/`cl100k_base`, ~4 bytes per token): https://github.com/openai/tiktoken
- Anthropic token counting docs (endpoint, estimate wording, unsupported inputs, free/rate limited, newer tokenizer note, caching FAQ): https://platform.claude.com/docs/en/build-with-claude/token-counting
- OpenAI cookbook, How to count tokens with tiktoken (fetched as the raw GitHub notebook; per-message and reply-priming overhead, estimate caveat): https://raw.githubusercontent.com/openai/openai-cookbook/main/examples/How_to_count_tokens_with_tiktoken.ipynb

Secondary — not re-fetched (hosts were egress-blocked this session):

- Gemini API token counting (countTokens): https://ai.google.dev/gemini-api/docs/tokens
- OpenAI cookbook rendered page for the same notebook: https://cookbook.openai.com/examples/how_to_count_tokens_with_tiktoken

---

## Related resources

- [Agent Cost and Latency Optimization](https://changegamer.ai/resources/agent-cost-latency-optimization.md): Practitioner reference for reducing the cost and latency of production AI agents: the compounding model, token-level levers (caching, pruning), request-level levers (Batch API, parallelism), model-level levers (routing, reasoning-effort controls), and architecture-level levers (step reduction, semantic caching, code offloading).
- [Agent Memory and Context Management](https://changegamer.ai/resources/agent-memory-context.md): Architecture reference for agent memory: types (working, long-term, episodic, semantic, procedural), context-management techniques (summarization, RAG, sliding windows, prompt caching), storage substrates, and memory frameworks — with security notes and cross-links to related guides.
- [Application-Level Response Caching for AI Agents](https://changegamer.ai/resources/agent-response-caching.md): How to implement exact-match and semantic caching in your agent application to eliminate redundant LLM calls, with threshold guidance, invalidation strategies, and a decision matrix for when semantic caching is unsafe.
- [How to Choose an LLM for Agentic Tasks](https://changegamer.ai/resources/choosing-an-llm-for-agents.md): A criteria-based decision framework for selecting an LLM for agent use: tool-calling reliability, long-context behavior, structured output, cost per task, latency, and a step-by-step selection procedure.

---

## Further reading

- [AI Agent Observability and the Production Evaluation Playbook](https://changegamer.ai/articles/agent-observability-and-evaluation.md): AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.
- [How to Track AI Agent Costs in Production](https://changegamer.ai/articles/agent-cost-telemetry-in-production.md): How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.

---

Index of all resources: https://changegamer.ai/llms.txt · Full corpus: https://changegamer.ai/llms-full.txt · Corpus data (NDJSON): https://changegamer.ai/api/corpus.jsonl · Offers: https://changegamer.ai/api/pricing.json
License the full corpus for RAG / fine-tuning (AI-use grant): https://changegamer.ai/corpus-license
