Understanding Tokens: The Building Blocks of LLMs
When you send a message to ChatGPT or Claude, the model doesn’t read your words the way you do. It breaks your text into tokens — chunks that might be whole words, parts of words, or even individual characters.
Understanding tokenization is surprisingly practical: it affects your API costs, your context window budget, and sometimes even the quality of your outputs.
What’s a Token?
A token is the smallest unit of text a language model processes. The word “understanding” might be a single token, while “tokenization” could be split into “token” and “ization.”
Here’s the rough math: 1 token ≈ 4 characters in English, or about ¾ of a word.
Why Should You Care?
Three reasons:
- Cost — API pricing is per-token. Knowing your token count helps you estimate costs.
- Context windows — Models have a maximum context measured in tokens. Your prompt + output must fit.
- Non-English text — Tokenizers trained primarily on English are less efficient with other languages.
Try It Yourself
The interactive visualizer below uses the same tokenizer (tiktoken) that GPT-4 uses. Type any text to see how it gets broken into tokens:
Interactive token visualizer will be embedded here in the full version.
Practical Tips
Want to go deeper? Check out the full lesson on tokens in our course.