LLM Tokenization
Large language models don't read characters or words; they read tokens. A tokenizer splits text into chunks from a fixed vocabulary, usually tens of thousands to a couple hundred thousand entries, and maps each chunk to an integer ID. The model only ever sees and predicts those IDs. Common words are often a single token, rarer words are split into pieces, and every character of every language can still be represented by falling back to bytes.
Tokenization sounds like plumbing, but it shapes everything practical about working with LLMs: what you pay (APIs bill per token), how much fits in a context window, how fast responses stream, and a long list of odd behaviors, from miscounting letters in a word to worse performance in some languages.
TL;DR
- A token is a chunk of text from the model's vocabulary: roughly ~4 characters or ~¾ of a word of English on average.
- Modern LLMs use subword tokenization, typically byte-level BPE (GPT, Llama, Claude-style models) or SentencePiece (Gemma, T5).
- Tokenizers are model-specific. The same text has different token counts on different models.
- Special tokens mark roles, message boundaries, and end-of-sequence; chat templates insert them.
- Count tokens with the provider's token-counting API or the model's tokenizer library before sending large prompts.
- Tokenization explains quirks such as weak character-level tasks, arithmetic on long numbers, and higher costs for non-English text and code.
Quick Example
Inspecting tokens with an open tokenizer (tiktoken, used by OpenAI models):
And counting tokens for a Claude request before sending it:
Notice that leading spaces belong to tokens (' isn', ' but') and that "Tokenization" becomes two pieces.
Core Concepts
Why Subwords?
- Word-level vocabularies can't handle unseen words, typos, or new names, and they'd need millions of entries.
- Character-level sequences are very long, which makes attention expensive and learning harder.
- Subword tokenization is the compromise: frequent words stay whole, rare words decompose into meaningful pieces (
unbelievable→un,believ,able), and any string can be encoded.
Byte-Pair Encoding (BPE)
BPE builds a vocabulary by starting from single bytes and repeatedly merging the most frequent adjacent pair in a training corpus into a new token, until the vocabulary reaches its target size. Encoding new text applies the learned merges. Byte-level BPE starts from the 256 possible bytes, so any UTF-8 text, emoji, or binary-looking string is representable, and nothing is ever "unknown".
WordPiece and SentencePiece / Unigram
- WordPiece (BERT) is similar to BPE but chooses merges that maximize training-data likelihood; continuation pieces are marked with
##. - SentencePiece treats input as a raw character stream, with spaces encoded as
▁, so it works for languages without spaces (Japanese, Chinese, Thai). It supports BPE or a Unigram language model that starts large and prunes.
Vocabulary Size
Vocabularies have grown from ~32k (early Llama) to ~100k–260k in current models. Larger vocabularies encode text in fewer tokens, especially multilingual text and code, which speeds generation and stretches the context window, at the cost of a larger embedding table.
Special Tokens and Chat Templates
Tokenizers reserve special tokens for structure: beginning and end of sequence, message and role boundaries, tool calls, and padding. Chat models are trained on conversations formatted with a specific chat template. Hosted APIs apply it for you. With open models, use the tokenizer's apply_chat_template rather than hand-formatting prompts, since a wrong template noticeably degrades output.
Tokens in Practice
Cost and Limits
API pricing is per million input and output tokens (output is typically several times more expensive), and context windows and max-output limits are measured in tokens. Rough English heuristics:
Code, JSON, non-Latin scripts, and heavy whitespace use more tokens per character. Images, audio, and PDFs are also converted into tokens by multimodal models.
Counting Tokens
- Use the provider's counting endpoint (Anthropic's
count_tokens, Gemini'scountTokens) or the model's official tokenizer (tiktoken, Hugging FaceAutoTokenizer). - Don't estimate one model's usage with another model's tokenizer; counts can differ by 10–30% or more.
- Read the
usagefield in API responses for what was actually billed, including cached input tokens (see LLM inference for prompt caching).
Quirks Tokenization Explains
- Spelling and counting letters: the model sees
strawberryas a couple of tokens, not ten letters, so character-level questions are hard without reasoning step by step. - Arithmetic on long numbers: digits may be grouped inconsistently into tokens; many newer tokenizers split numbers into single digits or fixed groups to help.
- Leading spaces matter:
" Paris"and"Paris"are different tokens. It's relevant when constraining or parsing outputs. - Language cost disparity: languages underrepresented in tokenizer training need more tokens per sentence, so they cost more and fit less in context.
- Glitch tokens: rare tokens seen during tokenizer training but barely during model training can produce bizarre outputs.
Best Practices
Budget in Tokens, Not Characters
Set limits (maximum document size, chat history length, retrieved chunks for RAG) in tokens measured with the target model's tokenizer. Character-based truncation can cut too much or overflow the context.
Chunk Documents on Semantic Boundaries
When splitting for embeddings or retrieval, target a token size (for example 300–800 tokens) but split on headings, paragraphs, or sentences, not mid-word, and add small overlaps. See embeddings.
Keep Prompts Lean
Verbose system prompts, repeated instructions, and pretty-printed JSON all cost tokens on every request. Compact formats (minified JSON, concise instructions) and prompt caching cut costs without changing behavior.
Use Structured Outputs Instead of Token Tricks
Rather than relying on logit biases or stop sequences keyed to specific tokens, use structured outputs or tool calling for reliable formats. They're robust to tokenizer differences between models.
Common Mistakes
Estimating With the Wrong Tokenizer
Truncating Strings by Characters
Cutting user input at 10,000 characters might be 2,500 tokens of English or 6,000 tokens of code, and can split a multi-byte character. Truncate by tokens using the tokenizer.
Forgetting Output Tokens
Context limits apply to input plus output. A prompt that fills the window leaves no room for the answer, and a low max_tokens truncates responses mid-sentence (check the stop reason).
FAQ
How many tokens is a word?
In English, about 1.3 tokens per word on average, or roughly 4 characters per token. Common words are usually one token, long or rare words several. Other languages and code vary widely, so measure with the model's tokenizer when it matters.
Why do different models count the same text differently?
Each model family trains its own tokenizer with its own vocabulary and merges. A larger or more multilingual vocabulary encodes the same text in fewer tokens. Even versions within one family can change tokenizers.
Do images and audio use tokens?
Yes. Multimodal models convert images, audio, and documents into token sequences: images typically cost hundreds to a few thousand tokens depending on resolution, and PDFs cost text tokens plus image tokens per page. Providers document their formulas and include them in usage reporting.
Will tokenization go away?
Research on byte-level and "tokenizer-free" models continues, and some architectures learn dynamic chunking of bytes. For now, production LLMs rely on subword tokenizers because shorter sequences make training and inference much cheaper.
Related Topics
- Large Language Models — How LLMs work end to end
- Context Windows — How many tokens a model can consider
- Transformer Architecture — What happens to token IDs next
- LLM Sampling — How the next token is chosen
- Embeddings — Chunking text for retrieval
- NLP — Tokenization in classic NLP pipelines