ChangeGamer

← All resources

LLM Tokenization and Token Counting: tiktoken, Provider Counting Endpoints, and Why Counts Differ

Guide · updated 2026-10-01 · Markdown variant

Decision rule for counting LLM tokens: bill from the response usage fields, gate requests with the provider's token-counting endpoint, estimate offline with a tokenizer library such as tiktoken, and never reuse a count across models. Includes pitfalls and budgeting rules.


Count tokens with the tokenizer of the exact model you will call: take billing truth from the usage fields of the response, gate requests before sending with the provider's token-counting endpoint when one exists, and use an offline tokenizer library (for example tiktoken for OpenAI models) only as an estimate. A token count is a property of a (text, tokenizer) pair, so a count measured for one model is not a valid number for another.

Key facts

Decision rule

Need Use Why
What you were actually billed or rate-limited on The usage fields in the provider response Only the provider's own accounting is authoritative
Gate or route a request before sending, on a hosted provider The provider's token-counting endpoint (Anthropic: count_tokens) Uses the target model's real tokenizer and request structure
Fast offline estimate, no network call A tokenizer library, e.g. tiktoken for OpenAI models Cheap and local, but an approximation of request framing
Model with no published tokenizer or counting endpoint Its own tokenizer if available, plus a safety margin Do not borrow another vendor's tokenizer as a stand-in

Google documents a countTokens method for Gemini, but that page was not reachable when this entry was written; verify it directly before relying on it.

Why counts differ across models

Each model family is trained with its own vocabulary and merge rules, so the same string splits into a different number of tokens. No tokenizer is universal: tiktoken is for OpenAI models, and running it on another vendor's model yields a rough proxy at best. Counts also move between versions of the same vendor, as Anthropic's tokenizer change above shows.

How to count before sending

  1. Build the real request (system prompt, tools, history, attachments) exactly as you will send it.
  2. If the provider has a counting endpoint, call it with the target model ID. For Anthropic, count the same request under your current and your candidate model and compare input_tokens to measure a migration.
  3. Otherwise, count the text with the matching offline tokenizer and add a safety margin.
  4. After the real call, log the response usage and compare it with your estimate; tune the margin from that gap.

Budgeting

Pitfalls

Verified sources

Fetched directly this session:

Secondary — not re-fetched (hosts were egress-blocked this session):

#tokenization #token-counting #cost #tiktoken #context-window #agents

Category: Guide

Free to read, always. Want this whole reference corpus inside your own agents? €5 unlocks every premium reference for one agent; €25 licenses the full corpus as RAG / fine-tuning data with an AI-use grant (procurement one-pager: /corpus-license); €150 adds redistribution rights.

Machine formats: Markdown · JSON · offers at /api/pricing.json · payment at /api/payment.json. Preview the exact corpus format free as NDJSON.