Transformers
The Transformer architecture, introduced in the 2017 paper “Attention Is All You Need,” is the foundation of virtually every modern AI model. GPT, Claude, Gemini, Llama — they’re all Transformers.
The Problem Transformers Solved
Before Transformers, the dominant architecture for language tasks was the Recurrent Neural Network (RNN). RNNs process text one word at a time, left to right, maintaining a hidden state.
The problem: RNNs forget. By the time they reach the end of a long sentence, the information from the beginning has faded. They also can’t be parallelized — each step depends on the previous one, making training slow.
Attention: The Key Insight
The Transformer’s breakthrough is the attention mechanism. Instead of processing words sequentially, attention lets the model look at all words simultaneously and decide which ones are relevant to each other.
When processing “The cat sat on the mat because it was tired,” attention lets the model connect “it” back to “cat” — even though they’re separated by several words.
Self-Attention Step by Step
For each word in a sequence, the model creates three vectors:
- Query (Q): “What am I looking for?”
- Key (K): “What do I contain?”
- Value (V): “What information do I carry?”
The attention score between two words is the dot product of one word’s Query with another word’s Key. High score = high relevance. These scores determine how much each word’s Value contributes to the output.
Attention(Q, K, V) = softmax(QKᵀ / √dₖ) · V
The √dₖ scaling prevents the dot products from getting too large, which would make the softmax function produce extremely peaked distributions.
Multi-Head Attention
A single attention mechanism captures one type of relationship. Multi-head attention runs multiple attention mechanisms in parallel — each “head” can learn different relationship types:
- One head might focus on syntactic relationships (subject-verb)
- Another on semantic relationships (synonyms, antonyms)
- Another on positional relationships (nearby words)
The Transformer Block
A complete Transformer block consists of:
- Multi-head self-attention: Look at all positions, find relevant context
- Add & normalize: Residual connection + layer normalization
- Feed-forward network: Process each position independently
- Add & normalize: Another residual connection
Modern LLMs stack dozens to over a hundred of these blocks.
Encoder vs. Decoder
The original Transformer has two halves:
- Encoder: Processes the full input and creates representations. Used in BERT, sentence embeddings.
- Decoder: Generates output one token at a time, attending to both the input and previously generated tokens. Used in GPT, Claude.
Most modern LLMs are decoder-only — they generate text autoregressively, one token at a time.
Why Transformers Won
- Parallelizable: Unlike RNNs, all positions can be processed simultaneously during training
- Long-range dependencies: Attention connects any two positions directly, regardless of distance
- Scalable: Performance improves predictably with more parameters, more data, and more compute
This scalability is why we’ve seen the explosive growth in model size — Transformers reward scale in a way previous architectures didn’t.
Think Like a Transformer
Now that you know how attention and next-token prediction work, try it yourself. In this game, you’ll see a sentence with a blank and four candidate tokens — rank them from most to least likely, just like a Transformer would.
Next, we’ll look at what the Transformer actually processes: tokens.