LLM Tokenization and Token Counting: tiktoken, Provider Counting Endpoints, and Why Counts Differ
Decision rule for counting LLM tokens: bill from the response usage fields, gate requests with the provider's token-counting endpoint, estimate offline with a tokenizer library such as tiktoken, and never reuse a count across models. Includes pitfalls and budgeting rules.
Count tokens with the tokenizer of the exact model you will call: take billing truth from the usage fields of the response, gate requests before sending with the provider's token-counting endpoint when one exists, and use an offline tokenizer library (for example tiktoken for OpenAI models) only as an estimate. A token count is a property of a (text, tokenizer) pair, so a count measured for one model is not a valid number for another.
Key facts
- tiktoken is described in its README as "a fast BPE tokeniser for use with OpenAI's models."
tiktoken.encoding_for_model("gpt-4o")returns the encoding for a model;o200k_baseandcl100k_baseare two of the encodings it ships. The README notes BPE is reversible and lossless and works on arbitrary text, and that a token averages about 4 bytes in practice. - Anthropic offers a token-counting endpoint (
POST /v1/messages/count_tokens) that accepts the same structure as a message request, including system prompts, tools, images and PDFs, and returnsinput_tokens. Its docs call the result an estimate that can differ "by a small amount" from actual usage, and say you are not billed for any system-added tokens it includes. - The endpoint is free to use but rate limited separately from message creation. It rejects some inputs the Messages API accepts: most server tools (web search, code execution and others), the MCP connector, and image or document blocks with a
urlorfilesource (send base64 instead). - Anthropic documents that models from Claude Opus 4.7 onward use a newer tokenizer that produces roughly 30 percent more tokens for the same text (content-dependent) and tells you to recount against the model you plan to use.
- OpenAI's cookbook chat-counting example adds 3 tokens per message and 3 tokens to prime the reply, and warns the count "may change from model to model"; treat it as "an estimate, not a timeless guarantee."
Decision rule
| Need | Use | Why |
|---|---|---|
| What you were actually billed or rate-limited on | The usage fields in the provider response |
Only the provider's own accounting is authoritative |
| Gate or route a request before sending, on a hosted provider | The provider's token-counting endpoint (Anthropic: count_tokens) |
Uses the target model's real tokenizer and request structure |
| Fast offline estimate, no network call | A tokenizer library, e.g. tiktoken for OpenAI models | Cheap and local, but an approximation of request framing |
| Model with no published tokenizer or counting endpoint | Its own tokenizer if available, plus a safety margin | Do not borrow another vendor's tokenizer as a stand-in |
Google documents a countTokens method for Gemini, but that page was not reachable when this entry was written; verify it directly before relying on it.
Why counts differ across models
Each model family is trained with its own vocabulary and merge rules, so the same string splits into a different number of tokens. No tokenizer is universal: tiktoken is for OpenAI models, and running it on another vendor's model yields a rough proxy at best. Counts also move between versions of the same vendor, as Anthropic's tokenizer change above shows.
How to count before sending
- Build the real request (system prompt, tools, history, attachments) exactly as you will send it.
- If the provider has a counting endpoint, call it with the target model ID. For Anthropic, count the same request under your current and your candidate model and compare
input_tokensto measure a migration. - Otherwise, count the text with the matching offline tokenizer and add a safety margin.
- After the real call, log the response
usageand compare it with your estimate; tune the margin from that gap.
Budgeting
- Input and output tokens are metered separately and priced separately by providers; budget them as two numbers, and reserve output headroom inside the context window. See /resources/context-window-management.
- Cached input is accounted differently from fresh input. Anthropic notes its counting endpoint does not apply caching logic, so it will not show cache savings; measure those from the real response. See /resources/prompt-caching-for-agents.
- Reasoning/thinking tokens can count toward input on later turns depending on the model; Anthropic documents this per model, so check before assuming history is free.
- For cost levers beyond counting, see /resources/agent-cost-latency-optimization. For picking a model by capability and context size, see /resources/choosing-an-llm-for-agents.
Pitfalls
- Estimate vs billed. Offline and endpoint counts are estimates; only response usage is billing truth.
- Multimodal tokens. Images and PDFs consume tokens by provider-specific rules; count them through the provider rather than guessing. See /resources/multimodal-agents.
- Tool schemas. Tool definitions are part of the input and are counted; a text-only count that omits them under-reports. See /resources/reliable-tool-calling and /resources/structured-outputs-and-json-mode.
- Tokenizer changes between model versions. Recount whenever you change model ID, even within one vendor.
- Message-framing overhead. Chat formats add tokens around each message beyond the visible text.
Verified sources
Fetched directly this session:
- tiktoken README (BPE tokeniser for OpenAI models,
encoding_for_model,o200k_base/cl100k_base, ~4 bytes per token): https://github.com/openai/tiktoken - Anthropic token counting docs (endpoint, estimate wording, unsupported inputs, free/rate limited, newer tokenizer note, caching FAQ): https://platform.claude.com/docs/en/build-with-claude/token-counting
- OpenAI cookbook, How to count tokens with tiktoken (fetched as the raw GitHub notebook; per-message and reply-priming overhead, estimate caveat): https://raw.githubusercontent.com/openai/openai-cookbook/main/examples/How_to_count_tokens_with_tiktoken.ipynb
Secondary — not re-fetched (hosts were egress-blocked this session):
- Gemini API token counting (countTokens): https://ai.google.dev/gemini-api/docs/tokens
- OpenAI cookbook rendered page for the same notebook: https://cookbook.openai.com/examples/how_to_count_tokens_with_tiktoken
Free to read, always. Want this whole reference corpus inside your own agents? €5 unlocks every premium reference for one agent; €25 licenses the full corpus as RAG / fine-tuning data with an AI-use grant (procurement one-pager: /corpus-license); €150 adds redistribution rights.