Retrieval Augmented Generation
RAG is a system pattern for giving an LLM the right context before it answers. Rather than expensive fine-tuning, RAG systems retrieve relevant context at query time — enabling cost efficiency, freshness, and scalability with terabytes of external knowledge.
Production-grade system anatomy
A production RAG system comprises interconnected stages:
- Ingestion pipeline — data loading with 160+ connectors, cleaning, deduplication, metadata extraction across unstructured text, semi-structured data, images, audio, and video
- Chunking strategy — semantic chunking (respects meaningful boundaries) beats fixed-size. Hierarchical chunking with contextual headers preserves location context. Optimal range: 256–1024 tokens
- Embedding and vector search — dense embeddings for semantic similarity, hybrid search combining dense + sparse (BM25) for recall + precision
- Retrieval — from basic vector search → top-k, to hybrid dense + keyword search, to advanced multi-query expansion and corrective RAG
- Generation — strategic context ordering, citation and attribution, constraint-based generation to prevent hallucinations
Context windows: RAG vs. long-context
| Aspect | RAG | Long-Context (CAG) |
|---|---|---|
| Knowledge base | Unlimited (terabytes) | Window-bounded |
| Cost | Pay only for relevant chunks | Pay for all tokens |
| Latency | Fast (small context) | Slower (full documents) |
| Reasoning | Susceptible to retrieval errors | Holistic understanding |
| Architecture | Complex (vector DB, embeddings) | Simple (direct context) |
Key libraries and frameworks
Orchestration: LangChain (TS/Python) for general-purpose workflows, LlamaIndex (TS/Python) for focused RAG with 160+ data connectors.
Vector databases: Pinecone (managed, easiest ops), Weaviate (GraphRAG, self-hosted), Qdrant (hybrid retrieval, compact), Milvus (open-source, Kubernetes-native).
Embedding models: Open-source (all-MiniLM-L6-v2, BGE-large) and API-based (OpenAI text-embedding-3-large).
Reranking: Cohere Rerank, ColBERT, SBERT cross-encoders.