Memory and Context Windows
Overview
Context window is the working memory an AI uses to answer you. It’s everything the model can “see” at once: your prompt, the back-and-forth so far, and any text you pasted. It’s measured in tokens, which are roughly chunks of words. When a chat or a document runs longer than that window can hold, the model starts losing track of the earlier parts. That’s why your AI seems to forget.
Context is the new scaling problem. The model needs curated, relevant business knowledge at request time—not a generic “search over all docs.”
What is the “lost in the middle” effect?
Where you put information inside the window matters as much as whether it fits. In a 2023 study, models did best when the key information sat at the very beginning or the very end of the input, and performance dropped noticeably when that same information was buried in the middle of a long context (Liu et al., Stanford / TACL, 2023). The researchers named it “lost in the middle.”
You understand how models work. Now let’s look at the practical infrastructure you’ll use to integrate them into web applications.
API-Based Access
Most AI applications use hosted APIs:
const response = await fetch('https://api.anthropic.com/v1/messages', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'x-api-key': process.env.ANTHROPIC_API_KEY,
'anthropic-version': '2023-06-01',
},
body: JSON.stringify({
model: 'claude-sonnet-4-20250514',
max_tokens: 1024,
messages: [{ role: 'user', content: 'Hello!' }],
}),
});
This is the pattern: HTTP request with your API key, model selection, and message payload.
SDKs
In practice, you’ll use official SDKs instead of raw fetch:
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
const response = await client.messages.create({
model: 'claude-sonnet-4-20250514',
max_tokens: 1024,
messages: [{ role: 'user', content: 'Hello!' }],
});
SDKs handle retry logic, streaming, type safety, and API versioning.
Rate Limits and Quotas
Every API has limits:
- Requests per minute (RPM): How many API calls you can make
- Tokens per minute (TPM): Total tokens processed per minute
- Tokens per day (TPD): Daily token budget
Your application needs to handle rate limit errors gracefully — typically with exponential backoff and queuing.
Cost Management
AI API costs can surprise you:
| Scenario | ~Cost |
|---|---|
| 1 chatbot message | 0.001–0.01 |
| Summarize a 10-page document | 0.01–0.05 |
| Process 10,000 customer reviews | 5–50 |
| RAG pipeline with 100K daily queries | 100–1,000/day |
Cost control strategies:
- Use the smallest model that works
- Cache responses for identical or similar inputs
- Set
max_tokensappropriately — don’t leave it at the maximum - Monitor usage with provider dashboards
Architecture Patterns
AI features in web applications typically follow one of these patterns:
- Server-side proxy: Your backend calls the AI API. The user’s browser never sees the API key.
- Edge function: AI calls from Cloudflare Workers, Vercel Edge, or similar. Low latency, runs close to users.
- Client-side with BYOK: Users provide their own API keys. Your server proxies the request for CORS. (This is what our course playground uses.)
- Background job: Long AI tasks run asynchronously with results stored and retrieved later.
We’ll implement several of these patterns in the course.
What’s Next
In the final lesson of this module, we’ll look at the limitations and risks of generative AI — what these models can’t do, and what you need to watch out for.