Tokens & Tokenization
Overview
LLMs break text into small, model-friendly units called tokens, created by a tokenizer; tokens are what get embedded, predicted, billed, and constrained by context windows, because this gives frontier models a flexible, efficient way to represent language across many scripts and domains.
What is a token?
- A token is the basic unit of text that a large language model processes: it can be a whole word, part of a word, punctuation, or even a spacing pattern.
- Modern LLMs use subword tokens, so a word like “techno-optimism” might be split into pieces such as
"techno","-","optim","ism"rather than treated as a single atom. - For English, a common approximation is that 1 token is about 4 characters or about three‑quarters of a word (so roughly 100 tokens ≈ 75 words), though actual counts depend on punctuation, language, and formatting.
Internally, each token is mapped to an integer ID (e.g., heard → 2, dog → 4), and the model sees text as a sequence of token IDs like [1, 2, 3, 4, 5, 6, 7, 3, 8] rather than human-readable strings.
Input tokens are the tokens that come into the model from your app: system prompt, user messages, tool definitions, retrieved RAG context, etc
Output tokens are the tokens in the visible response the model sends back—what you actually display in your UI or consume in your code
Reasoning tokens (OpenAI) or thinking tokens (Anthropic, DeepSeek) are tokens the model uses for internal chains of thought—planning, intermediate steps, decomposing the problem—before it writes the final answer.
End-of-Sequence (EOS) Tokens: Special learned tokens (like
<|endoftext|>or</s>) generated naturally when the model determines an answer is complete.
Why LLMs use tokens instead of words
There are deep technical reasons why frontier LLMs are built around tokens rather than human words:
- Manageable vocabulary size
- A pure word-level model would need a vocabulary containing every possible word across all languages, slang, product names, code identifiers, etc., which is effectively unbounded.
- Subword tokenization lets the model have a fixed vocabulary (tens of thousands of tokens) that can still represent arbitrarily complex or new words by composing smaller units.
- Handling rare and out-of-vocabulary words
- Word-level models struggle when they encounter a word not present in their vocabulary; it becomes an “unknown” token with no internal structure.
- Subword tokens allow the model to represent rare words as sequences of known pieces, preserving morphology and allowing generalization (e.g., if it understands “optim” and “ism”, it can reason about “techno-optimism” even if that exact string was rare in training).
- Multilingual and noisy text robustness
- Character-level or subword-level tokenization handles different scripts (Latin, Devanagari, Cyrillic, etc.), emojis, URLs, and code without needing handcrafted word rules for each language.
- This makes frontier models more robust to typos, spacing quirks, and user-generated content, which are common in real-world inputs.
- Efficient training with embeddings and transformers
- Each token ID is associated with an embedding vector capturing its semantic relationships—how often it appears with other tokens or in similar contexts.
- Transformers operate over sequences of these token embeddings, predicting a probability distribution over the next token at each step; this is simpler and more efficient than trying to learn over raw characters or undefined “word” units with ambiguous boundaries.
- Direct alignment with pricing and limits
- Because the same unit (token) drives both model internals and API economics—embeddings, attention, context size, and per-request billing—providers and engineers can reason consistently about capacity and cost.
What is a tokenizer?
https://platform.openai.com/tokenizer Tokenizer
- A tokenizer is the algorithm (plus vocabulary) that converts raw text into tokens and back again.
- During training and inference, tokenization is the first step: text is split into tokens, each token is looked up in a vocabulary, and its ID and embedding are fed into the transformer.
- Different models use different tokenization schemes; common ones include Byte-Pair Encoding (BPE), WordPiece, and Unigram tokenization, all of which produce subword tokens from character sequences.
Frontier models like the GPT series use BPE-style subword tokenization, implemented in libraries such as tiktoken, to balance vocabulary size, multilingual coverage, and robustness to typos and novel words
A tokenizer may assign different numerical representations for the same word depending on its meaning in context. For example, the word “lie” could refer to a resting position or to saying something untruthful.
Tokenization
BPE starts with individual characters and iteratively merges the most frequent adjacent pairs to create longer subwords, forming a deterministic vocabulary. This is the encoding used by OpenAI for their ChatGPT models.
SentencePiece operates on raw text, supports a probabilistic Unigram model and a BPE variant, and better handles multilingual or noise‑rich corpora.
How Tokenization Works
Most modern models use Byte Pair Encoding (BPE):
- Start with individual characters as tokens
- Find the most frequent pair of adjacent tokens
- Merge them into a new token
- Repeat until you reach the desired vocabulary size
This creates a vocabulary where:
- Common words are single tokens (“the”, “is”, “and”)
- Less common words are split into sub-word tokens
- Rare words are broken into many small tokens
Why Tokenization Matters
Cost
API pricing is per-token. A 1,000-word prompt is roughly 1,300-1,500 tokens. Understanding this helps you estimate costs and optimize prompts.
Context Windows
Models have a maximum context window measured in tokens (e.g., 128K tokens). This is the total budget for your prompt plus the response. Longer prompts leave less room for output.
Different Tokenizers
Each model family uses its own tokenizer. The same text produces different token counts across models:
- GPT-4: uses
cl100k_basetokenizer - Claude: uses its own tokenizer
- Llama: uses SentencePiece
This means switching models can change your cost structure.
Non-English Text
Tokenizers trained primarily on English text are less efficient with other languages. Chinese, Japanese, and Arabic text often produces more tokens per character, increasing costs.
Code
Code is tokenized too, but programming language syntax maps to tokens differently than natural language. Heavily formatted or commented code uses more tokens.
BPE and SentencePiece are two major subword tokenization approaches, but they differ in flexibility. The benefit of BPE is its simplicity and speed, as it repeatedly merges frequent character pairs. SentencePiece operates directly on raw text, making it more suitable for multilingual or noisy corpora.
Token economics (cost, limits, engineering)
From the perspective of model providers, the faster tokens can be processed, the faster models can learn and respond. The goal is to achieve the fastest processing time and lowest cost per token to optimize AI infrastructure and maximize revenue generation.
From an engineering and business perspective, “token economics” is about how tokens determine cost, speed, and scalability of your Gen AI system. Tokens are the “machine-level” units of language: they are what you must budget, measure, and optimize when you design prompts, chunk documents for RAG, or scale an LLM-backed application.
Cached input tokens: tokens in a reusable prefix (system prompt, static instructions, etc.) that benefit from provider‑side prompt caching, and are billed at a discounted “cached input” rate.
LLMs don’t read words. They read tokens — chunks of text that might be whole words, parts of words, or even individual characters.
Understanding tokenization is practical knowledge: it directly affects your API costs, context window usage, and output quality.
Rough rule of thumb: 1 token ≈ 4 characters in English, or about ¾ of a word.
Working With Token Limits
When building applications:
- Count tokens before sending — don’t guess
- Reserve space for output — if context is 128K and your prompt is 120K, the model can only generate 8K tokens of output
- Truncate strategically — when you must fit content into a token budget, cut less important content, not random chunks
- Use tiktoken — OpenAI’s open-source tokenizer library works in JavaScript (via Wasm)
In the next lesson, we’ll look at how models actually generate text from tokens.