Retrieval Augmented Generation

14 min read

RAG is a system pattern for giving an LLM the right context before it answers. Rather than expensive fine-tuning, RAG systems retrieve relevant context at query time — enabling cost efficiency, freshness, and scalability with terabytes of external knowledge.

Production-grade system anatomy

A production RAG system comprises interconnected stages:

  1. Ingestion pipeline — data loading with 160+ connectors, cleaning, deduplication, metadata extraction across unstructured text, semi-structured data, images, audio, and video
  2. Chunking strategy — semantic chunking (respects meaningful boundaries) beats fixed-size. Hierarchical chunking with contextual headers preserves location context. Optimal range: 256–1024 tokens
  3. Embedding and vector search — dense embeddings for semantic similarity, hybrid search combining dense + sparse (BM25) for recall + precision
  4. Retrieval — from basic vector search → top-k, to hybrid dense + keyword search, to advanced multi-query expansion and corrective RAG
  5. Generation — strategic context ordering, citation and attribution, constraint-based generation to prevent hallucinations

Context windows: RAG vs. long-context

Aspect RAG Long-Context (CAG)
Knowledge base Unlimited (terabytes) Window-bounded
Cost Pay only for relevant chunks Pay for all tokens
Latency Fast (small context) Slower (full documents)
Reasoning Susceptible to retrieval errors Holistic understanding
Architecture Complex (vector DB, embeddings) Simple (direct context)

Key libraries and frameworks

Orchestration: LangChain (TS/Python) for general-purpose workflows, LlamaIndex (TS/Python) for focused RAG with 160+ data connectors.

Vector databases: Pinecone (managed, easiest ops), Weaviate (GraphRAG, self-hosted), Qdrant (hybrid retrieval, compact), Milvus (open-source, Kubernetes-native).

Embedding models: Open-source (all-MiniLM-L6-v2, BGE-large) and API-based (OpenAI text-embedding-3-large).

Reranking: Cohere Rerank, ColBERT, SBERT cross-encoders.

Further reading