Inference
Overview
Inference is the process where a trained machine learning model evaluates new input data and generates outputs—predictions, classifications, or generated text. The model applies learned patterns from its training phase to produce results for production use cases.
When we type a question into a chatbot, the response feels instant. During inference, an AI receives a prompt — which, depending on the model, may be text, image, audio clip, video, sensor data or even gene sequence — that it translates into a series of tokens. The model processes these input tokens, generates its response as tokens and then translates it to the user’s expected format.
Model training—which happens in data science labs with large datasets and iterative experimentation uses supervised or unsupervised learning to teach the model data patterns.
Inference happens across three conceptual steps: the input data is preprocessed, fed through the model’s learned parameters (weights and layers), and then the output is post-processed into the expected format.
Infrastructure
vLLM
vLLM is the current industry standard for high-throughput LLM serving.
Performance Metrics
Time to First Token (TTFT)
This metric shows how long a user needs to wait before seeing the model’s output.
The latency between a user submitting a prompt and the AI model starting to respond, and inter-token or token-to-token latency, the rate at which subsequent output tokens are generated, determine how an end user experiences the output of an AI application. The longer the prompt, the larger the TTFT. This is because the attention mechanism requires the whole input sequence to compute and create the so-called key-value cache.
Tokens Per Second (TPS)
Total TPS per system represents the total output tokens per seconds throughput, accounting for all the requests happening simultaneously. As the number of requests increases, the total TPS per system increases, until it reaches a saturation point for all the available GPU compute resources, beyond which it might decrease.
Inter-token Latency (ITL)
This is defined as the average time between consecutive tokens and is also known as time per output token (TPOT)
Inference Costs
GPUs optimized for inference (A100, H100, H200) are scarce and expensive.
Optimizations
The KV cache is primarily an inference-time optimization. During inference, the KV cache stores key and value tensors for previously processed tokens, so the model can generate new tokens efficiently without recomputing attention for the whole sequence.
Training teaches the model. Inference is when the model actually generates text. Understanding how inference works helps you control output quality through parameters like temperature, top-k, and top-p.
Autoregressive Generation
LLMs generate text one token at a time. For each position, the model:
- Takes all previous tokens as input
- Processes them through the Transformer layers
- Produces a probability distribution over the entire vocabulary
- Selects one token from that distribution
- Appends it to the sequence and repeats
This is autoregressive generation — each new token depends on all tokens before it.
The Probability Distribution
At each step, the model doesn’t output a single token — it outputs a probability for every token in its vocabulary (typically 32K–100K+ tokens).
For “The cat sat on the ___”:
- “mat” → 12%
- “floor” → 8%
- “couch” → 6%
- “ground” → 5%
- …thousands more tokens with smaller probabilities
How you select from this distribution determines the character of the output.
Sampling Parameters
Temperature
Controls randomness. Technically, it scales the logits (raw scores) before applying softmax.
- Temperature = 0: Always pick the highest-probability token (deterministic, repetitive)
- Temperature = 0.3–0.7: Balanced (good for most applications)
- Temperature = 1.0: Full diversity (creative, but less coherent)
- Temperature > 1.0: Very random (often incoherent)
Top-k
Only consider the top k most likely tokens. If k=50, the model ignores all but the 50 highest-probability tokens.
Top-p (Nucleus Sampling)
Only consider tokens whose cumulative probability reaches p. If p=0.9, include tokens until their combined probability hits 90%.
Top-p is generally preferred over top-k because it adapts to the distribution — sometimes only 5 tokens are needed to reach 90%, sometimes 500.
Practical Guidelines
| Use Case | Temperature | Top-p |
|---|---|---|
| Code generation | 0–0.2 | 0.95 |
| Factual Q&A | 0–0.3 | 0.9 |
| Creative writing | 0.7–1.0 | 0.95 |
| Brainstorming | 0.8–1.2 | 1.0 |
| Classification | 0 | 1.0 |
Latency and Streaming
Since tokens are generated sequentially, inference has inherent latency. Two key metrics:
- Time to First Token (TTFT): How long until the first token appears
- Tokens per Second (TPS): How fast subsequent tokens are generated
Streaming sends tokens to the client as they’re generated, rather than waiting for the complete response. This dramatically improves perceived performance — users see text appearing immediately instead of waiting.
We’ll implement streaming in Module 2 when we build our first AI-powered features.