Models
Overview
A Model is a simplified mathematical or computational representation of a real-world system, process, or relationship that is used to make predictions or understand patterns. In AI and machine learning, it’s specifically an algorithm that has been trained on data to perform tasks like classification, prediction, or generation.
In an AI application, the model is — the fundamental system component.
Categories of AI Models
These differ mainly in how they’re trained, what inputs they accept, how they behave at inference, and what applications they’re suited for.
There are several major “types” of modern AI/LLM models, each optimized for a different role:
Modern providers ship families of models tuned for different jobs:
- Base/Foundation models
- Deep research & “thinking” models (LLMs)
- Multi-modal models
- Agentic / tools‑first variants
- Cost‑effective models
- Code models specialized on programming languages and code repositories
- Embedding models are trained to map inputs (text, and sometimes images) into vector representations suitable for semantic search and RAG
- Small Language Models (SLMs) are compact LMs designed for low‑latency, resource‑constrained environments
- World models builds a latent representation of “the world” (often spatial‑temporal scenes) and learns how that world changes over time given actions
Throughout the course, you will learn about various different models and how they facilitate building different applications.
Why this is important?
For AI engineering work, the key is to treat “model type” as a capability contract: pick reasoning models for hard problem‑solving; tool‑calling/chat models for interactive agents; multimodal models for anything that must see/hear; embedding models as infrastructure for retrieval; and base/instruct/code variants depending on how much control and specialization you need.
OpenAI’s o1, a step-by-step reasoning model, solved 83% of problems on the 2024 AIME math exam, against 13% for the non-reasoning GPT-4o (OpenAI o1, via Wikipedia, 2024
The AI model landscape changes fast. New models drop monthly, benchmarks shift, and yesterday’s state-of-the-art becomes today’s baseline. Here’s how to navigate it.
The Major Families
OpenAI (GPT Series)
- GPT-4o: Multimodal flagship. Strong at reasoning, coding, and following complex instructions.
- GPT-4o mini: Smaller, faster, cheaper. Good enough for most production use cases.
- o1/o3 series: Reasoning-focused models that “think” before answering.
Anthropic (Claude Series)
- Claude Opus: Most capable. Excels at nuanced analysis, long documents, and complex tasks.
- Claude Sonnet: Balanced speed and capability. Most popular for production.
- Claude Haiku: Fast and cheap. Great for high-volume, simpler tasks.
Google (Gemini Series)
- Gemini Ultra/Pro: Google’s flagship multimodal models.
- Gemini Flash: Optimized for speed, built for high-throughput applications.
Others
- Llama: Open-weight models you can run locally or fine-tune.
- Mistral: Strong open models from a French startup.
- Cohere: Enterprise-focused with strong RAG capabilities.
- Deepseek: Competitive open models at lower training costs.
How to Choose
The model choice depends on your use case:
| Factor | Small Model | Large Model |
|---|---|---|
| Latency | Fast | Slower |
| Cost | Cheap | Expensive |
| Quality | Good for simple tasks | Better for complex reasoning |
| Context | Shorter context windows | Longer context windows |
API vs. Self-Hosted
Most web developers will use API-based models (OpenAI, Anthropic, Google). Self-hosting eliminates API costs and data privacy concerns.
- Data privacy requirements prevent sending data to third parties
- You need to fine-tune for a specific domain
- API costs at your scale exceed infrastructure costs
- You need guaranteed availability
We’ll cover both approaches in this course.