{
  "slug": "llm-tokenization-and-token-counting",
  "title": "LLM Tokenization and Token Counting: tiktoken, Provider Counting Endpoints, and Why Counts Differ",
  "description": "Decision rule for counting LLM tokens: bill from the response usage fields, gate requests with the provider's token-counting endpoint, estimate offline with a tokenizer library such as tiktoken, and never reuse a count across models. Includes pitfalls and budgeting rules.",
  "category": "Guide",
  "tags": [
    "tokenization",
    "token-counting",
    "cost",
    "tiktoken",
    "context-window",
    "agents"
  ],
  "updated": "2026-10-01",
  "premium": false,
  "rights": {
    "access": "free",
    "license": "https://changegamer.ai/license.xml",
    "pricing": "https://changegamer.ai/api/pricing.json",
    "payment": "https://changegamer.ai/api/payment.json"
  },
  "canonical": "https://changegamer.ai/resources/llm-tokenization-and-token-counting",
  "markdown": "https://changegamer.ai/resources/llm-tokenization-and-token-counting.md",
  "outline": [
    {
      "depth": 2,
      "text": "Key facts",
      "anchor": "key-facts"
    },
    {
      "depth": 2,
      "text": "Decision rule",
      "anchor": "decision-rule"
    },
    {
      "depth": 2,
      "text": "Why counts differ across models",
      "anchor": "why-counts-differ-across-models"
    },
    {
      "depth": 2,
      "text": "How to count before sending",
      "anchor": "how-to-count-before-sending"
    },
    {
      "depth": 2,
      "text": "Budgeting",
      "anchor": "budgeting"
    },
    {
      "depth": 2,
      "text": "Pitfalls",
      "anchor": "pitfalls"
    },
    {
      "depth": 2,
      "text": "Verified sources",
      "anchor": "verified-sources"
    }
  ],
  "related": [
    {
      "slug": "agent-cost-latency-optimization",
      "title": "Agent Cost and Latency Optimization",
      "description": "Practitioner reference for reducing the cost and latency of production AI agents: the compounding model, token-level levers (caching, pruning), request-level levers (Batch API, parallelism), model-level levers (routing, reasoning-effort controls), and architecture-level levers (step reduction, semantic caching, code offloading).",
      "url": "https://changegamer.ai/resources/agent-cost-latency-optimization"
    },
    {
      "slug": "agent-memory-context",
      "title": "Agent Memory and Context Management",
      "description": "Architecture reference for agent memory: types (working, long-term, episodic, semantic, procedural), context-management techniques (summarization, RAG, sliding windows, prompt caching), storage substrates, and memory frameworks — with security notes and cross-links to related guides.",
      "url": "https://changegamer.ai/resources/agent-memory-context"
    },
    {
      "slug": "agent-response-caching",
      "title": "Application-Level Response Caching for AI Agents",
      "description": "How to implement exact-match and semantic caching in your agent application to eliminate redundant LLM calls, with threshold guidance, invalidation strategies, and a decision matrix for when semantic caching is unsafe.",
      "url": "https://changegamer.ai/resources/agent-response-caching"
    },
    {
      "slug": "choosing-an-llm-for-agents",
      "title": "How to Choose an LLM for Agentic Tasks",
      "description": "A criteria-based decision framework for selecting an LLM for agent use: tool-calling reliability, long-context behavior, structured output, cost per task, latency, and a step-by-step selection procedure.",
      "url": "https://changegamer.ai/resources/choosing-an-llm-for-agents"
    }
  ],
  "furtherReading": [
    {
      "slug": "agent-observability-and-evaluation",
      "title": "AI Agent Observability and the Production Evaluation Playbook",
      "description": "AI agent observability and evaluation once an agent is already live: tracing spans, retrieval-attribution logs, judge-based screening vs. human acceptance, online metrics, cost telemetry, and incident postmortems that feed evaluation fixtures.",
      "url": "https://changegamer.ai/articles/agent-observability-and-evaluation"
    },
    {
      "slug": "agent-cost-telemetry-in-production",
      "title": "How to Track AI Agent Costs in Production",
      "description": "How to track AI agent costs in production: version your price table, separate cached from uncached tokens, keep failed runs in the books, attribute shared costs, reconcile against the invoice, and alert on rate of spend.",
      "url": "https://changegamer.ai/articles/agent-cost-telemetry-in-production"
    }
  ],
  "body": "Count tokens with the tokenizer of the exact model you will call: take billing truth from the `usage` fields of the response, gate requests before sending with the provider's token-counting endpoint when one exists, and use an offline tokenizer library (for example tiktoken for OpenAI models) only as an estimate. A token count is a property of a (text, tokenizer) pair, so a count measured for one model is not a valid number for another.\n\n## Key facts\n\n- **tiktoken** is described in its README as \"a fast BPE tokeniser for use with OpenAI's models.\" `tiktoken.encoding_for_model(\"gpt-4o\")` returns the encoding for a model; `o200k_base` and `cl100k_base` are two of the encodings it ships. The README notes BPE is reversible and lossless and works on arbitrary text, and that a token averages about 4 bytes in practice.\n- **Anthropic** offers a token-counting endpoint (`POST /v1/messages/count_tokens`) that accepts the same structure as a message request, including system prompts, tools, images and PDFs, and returns `input_tokens`. Its docs call the result an **estimate** that can differ \"by a small amount\" from actual usage, and say you are not billed for any system-added tokens it includes.\n- The endpoint is **free to use** but rate limited separately from message creation. It rejects some inputs the Messages API accepts: most server tools (web search, code execution and others), the MCP connector, and image or document blocks with a `url` or `file` source (send base64 instead).\n- Anthropic documents that models from Claude Opus 4.7 onward use a newer tokenizer that produces roughly 30 percent more tokens for the same text (content-dependent) and tells you to recount against the model you plan to use.\n- OpenAI's cookbook chat-counting example adds 3 tokens per message and 3 tokens to prime the reply, and warns the count \"may change from model to model\"; treat it as \"an estimate, not a timeless guarantee.\"\n\n## Decision rule\n\n| Need | Use | Why |\n|---|---|---|\n| What you were actually billed or rate-limited on | The `usage` fields in the provider response | Only the provider's own accounting is authoritative |\n| Gate or route a request before sending, on a hosted provider | The provider's token-counting endpoint (Anthropic: `count_tokens`) | Uses the target model's real tokenizer and request structure |\n| Fast offline estimate, no network call | A tokenizer library, e.g. tiktoken for OpenAI models | Cheap and local, but an approximation of request framing |\n| Model with no published tokenizer or counting endpoint | Its own tokenizer if available, plus a safety margin | Do not borrow another vendor's tokenizer as a stand-in |\n\nGoogle documents a countTokens method for Gemini, but that page was not reachable when this entry was written; verify it directly before relying on it.\n\n## Why counts differ across models\n\nEach model family is trained with its own vocabulary and merge rules, so the same string splits into a different number of tokens. No tokenizer is universal: tiktoken is for OpenAI models, and running it on another vendor's model yields a rough proxy at best. Counts also move between versions of the same vendor, as Anthropic's tokenizer change above shows.\n\n## How to count before sending\n\n1. Build the real request (system prompt, tools, history, attachments) exactly as you will send it.\n2. If the provider has a counting endpoint, call it with the target model ID. For Anthropic, count the same request under your current and your candidate model and compare `input_tokens` to measure a migration.\n3. Otherwise, count the text with the matching offline tokenizer and add a safety margin.\n4. After the real call, log the response `usage` and compare it with your estimate; tune the margin from that gap.\n\n## Budgeting\n\n- Input and output tokens are metered separately and priced separately by providers; budget them as two numbers, and reserve output headroom inside the context window. See /resources/context-window-management.\n- Cached input is accounted differently from fresh input. Anthropic notes its counting endpoint does not apply caching logic, so it will not show cache savings; measure those from the real response. See /resources/prompt-caching-for-agents.\n- Reasoning/thinking tokens can count toward input on later turns depending on the model; Anthropic documents this per model, so check before assuming history is free.\n- For cost levers beyond counting, see /resources/agent-cost-latency-optimization. For picking a model by capability and context size, see /resources/choosing-an-llm-for-agents.\n\n## Pitfalls\n\n- **Estimate vs billed.** Offline and endpoint counts are estimates; only response usage is billing truth.\n- **Multimodal tokens.** Images and PDFs consume tokens by provider-specific rules; count them through the provider rather than guessing. See /resources/multimodal-agents.\n- **Tool schemas.** Tool definitions are part of the input and are counted; a text-only count that omits them under-reports. See /resources/reliable-tool-calling and /resources/structured-outputs-and-json-mode.\n- **Tokenizer changes between model versions.** Recount whenever you change model ID, even within one vendor.\n- **Message-framing overhead.** Chat formats add tokens around each message beyond the visible text.\n\n## Verified sources\n\nFetched directly this session:\n\n- tiktoken README (BPE tokeniser for OpenAI models, `encoding_for_model`, `o200k_base`/`cl100k_base`, ~4 bytes per token): https://github.com/openai/tiktoken\n- Anthropic token counting docs (endpoint, estimate wording, unsupported inputs, free/rate limited, newer tokenizer note, caching FAQ): https://platform.claude.com/docs/en/build-with-claude/token-counting\n- OpenAI cookbook, How to count tokens with tiktoken (fetched as the raw GitHub notebook; per-message and reply-priming overhead, estimate caveat): https://raw.githubusercontent.com/openai/openai-cookbook/main/examples/How_to_count_tokens_with_tiktoken.ipynb\n\nSecondary — not re-fetched (hosts were egress-blocked this session):\n\n- Gemini API token counting (countTokens): https://ai.google.dev/gemini-api/docs/tokens\n- OpenAI cookbook rendered page for the same notebook: https://cookbook.openai.com/examples/how_to_count_tokens_with_tiktoken",
  "sources": [
    "https://github.com/openai/tiktoken",
    "https://platform.claude.com/docs/en/build-with-claude/token-counting",
    "https://raw.githubusercontent.com/openai/openai-cookbook/main/examples/How_to_count_tokens_with_tiktoken.ipynb",
    "https://ai.google.dev/gemini-api/docs/tokens",
    "https://cookbook.openai.com/examples/how_to_count_tokens_with_tiktoken"
  ]
}