TempatGunting

Retrieval Augmented Generation: Architecture Patterns for Production RAG

Why Tutorial RAG Falls Apart in Production

Every RAG tutorial follows the same script: chunk your documents, embed them, store vectors, retrieve top-k, generate an answer. It works beautifully on the demo. Then you deploy it against a real corpus of 10 million documents with users who ask ambiguous questions, and the entire architecture collapses under its own assumptions.

Production RAG requires a fundamentally different approach. The retrieval pipeline alone needs three stages instead of one. The chunking strategy must understand document structure rather than splitting on arbitrary token counts. The generation layer needs guardrails that catch hallucination before the response reaches the user. And you need an evaluation framework that can distinguish between retrieval failures and generation failures — because the fixes are completely different.

After building RAG systems across legal document search, medical knowledge bases, and enterprise support portals, I've identified the architecture patterns that consistently survive contact with production traffic. These aren't theoretical improvements — they're the minimum viable architecture for RAG systems that handle more than a hundred queries per minute.

User Query raw input Query Transform expansion + rewrite BM25 Sparse keyword recall Dense Vector semantic recall RRF Merge + Re-rank LLM Generate + guardrails

Query Transformation: The First Stage Most Teams Skip

Raw user queries are terrible retrieval inputs. Users type "why isn't the thing working" when they mean "error handling for authentication token expiration in the payment service." The gap between user intent and retrieval-friendly queries is where most RAG systems lose accuracy.

The query transformation pipeline rewrites the raw input into one or more retrieval-optimized queries. This isn't prompt engineering — it's a dedicated model or LLM call that produces structured search queries from ambiguous natural language. The three transformation techniques that matter most are query expansion, hypothetical document embedding (HyDE), and step-back prompting.

Query expansion generates multiple related queries from the original input. A question about "PyTorch memory errors during training" might expand to queries about CUDA OOM handling, gradient accumulation as a memory reduction strategy, and mixed-precision training for memory efficiency. Each expanded query retrieves independently, and the results merge before re-ranking.

class QueryTransformer:
    def __init__(self, expansion_model, hyde_model):
        self.expansion = expansion_model
        self.hyde = hyde_model

    def transform(self, raw_query: str) -> list[str]:
        expanded = self.expansion.generate(
            f"Generate 3 search queries for: {raw_query}"
        )
        hypothetical_doc = self.hyde.generate(
            f"Write a paragraph answering: {raw_query}"
        )
        return [raw_query] + expanded + [hypothetical_doc]

HyDE takes a different approach: instead of transforming the query, it generates a hypothetical document that would answer the query, then uses that generated text as the retrieval input. The intuition is that document-to-document similarity in embedding space is more reliable than query-to-document similarity. In our benchmarks, HyDE improved recall@10 by 12% on technical documentation corpora.

Hybrid Retrieval: Sparse Plus Dense

Pure dense retrieval with a single vector database fails on two classes of queries: exact keyword matches (product IDs, error codes, API endpoint names) and rare terms that the embedding model hasn't seen enough times during training to encode meaningfully.

Hybrid retrieval combines BM25 sparse search with dense vector search. BM25 handles the exact-match cases perfectly — it finds the document containing error code "ERR_AUTH_0x4F" every time, while the embedding model might map that code to something semantically adjacent but factually wrong. Dense search handles natural language queries where the exact words don't appear in the target document.

The combination uses reciprocal rank fusion to merge results from both systems:

def reciprocal_rank_fusion(
    sparse_results: list, dense_results: list, k: int = 60
) -> list:
    scores = {}
    for rank, doc in enumerate(sparse_results):
        scores[doc.id] = scores.get(doc.id, 0) + 1.0 / (k + rank + 1)
    for rank, doc in enumerate(dense_results):
        scores[doc.id] = scores.get(doc.id, 0) + 1.0 / (k + rank + 1)
    return sorted(scores.items(), key=lambda x: -x[1])

In our production system, hybrid retrieval with RRF consistently outperforms either sparse or dense alone by 15-25% on recall@20. The improvement is most dramatic on mixed queries that combine technical terms with natural language context — exactly the type of queries real users submit.

Semantic Chunking: Structure-Aware Document Splitting

Fixed-size chunking with arbitrary token windows is the single most common mistake in production RAG. A 512-token chunk that splits a code example in half, or separates a table header from its data rows, produces embedding vectors that capture noise rather than meaning.

Semantic chunking respects the document's inherent structure. It identifies section boundaries from headings, paragraph breaks, code fences, and list markers. Each chunk represents a coherent unit of information. The chunking algorithm walks the document tree and creates chunks at natural break points, with overlap windows that include parent section context.

For hierarchical documents — technical manuals, legal contracts, API documentation — we add parent document references to each chunk. When a chunk is retrieved, the re-ranker can pull in the parent section for additional context. This hierarchical approach resolves the tension between small chunks (better retrieval precision) and large chunks (better generation context).

The approach connects to how we handle document chunking at scale and relates to broader challenges in embedding quality management.

Cross-Encoder Re-Ranking

The initial retrieval stage optimizes for recall — cast a wide net and don't miss relevant documents. Re-ranking optimizes for precision — sort the retrieved candidates by true relevance to the query. The difference in quality between bi-encoder retrieval scores and cross-encoder re-ranking scores is dramatic.

Bi-encoders compute query and document embeddings independently, then measure similarity with a dot product. This is fast but loses the interaction between query and document tokens. Cross-encoders process the query-document pair jointly through a transformer, capturing fine-grained relevance signals that bi-encoders miss.

The tradeoff is latency: cross-encoders are 100x slower than bi-encoder similarity search. We retrieve 100 candidates with bi-encoders in 50ms, then re-rank with a cross-encoder in 200-400ms. The re-ranked top-10 results contain the correct answer 85% of the time, compared to 62% for the raw bi-encoder top-10. This relates directly to how model serving optimization handles the latency-quality tradeoff.

from sentence_transformers import CrossEncoder

reranker = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-12-v2")

def rerank(query: str, candidates: list[dict], top_k: int = 10):
    pairs = [(query, doc["text"]) for doc in candidates]
    scores = reranker.predict(pairs)
    ranked = sorted(
        zip(candidates, scores), key=lambda x: -x[1]
    )
    return [doc for doc, score in ranked[:top_k]]

Hallucination Guardrails

Even with perfect retrieval, LLMs hallucinate. They'll confidently generate information that isn't in any retrieved document, mixing parametric knowledge from training with the provided context in ways that are impossible for users to detect. Production RAG needs active guardrails, not just hope.

Our guardrail pipeline runs three checks on every generated response. First, a natural language inference (NLI) model classifies each claim in the response as supported, contradicted, or neutral relative to the retrieved context. Claims classified as contradicted or neutral get flagged. Second, a citation verification step checks that every factual statement maps to a specific chunk in the retrieved context. Third, a confidence calibration module estimates the model's certainty and triggers a "I don't have enough information" response when confidence falls below threshold.

The NLI-based factuality check catches approximately 70% of hallucinations in our benchmarks. Combined with citation verification, the detection rate improves to 85%. The remaining 15% are subtle hallucinations where the model makes plausible inferences from the context that happen to be wrong — those require human review. For more on handling model outputs safely, see our guide on knowledge distillation for reliable inference.

LLM Response raw generation NLI Check claim verification Citation Match source mapping Confidence Gate threshold filter Safe Response Flagged / Refused

Evaluation Frameworks That Actually Work

End-to-end RAG evaluation is meaningless without component-level metrics. When the system gives a wrong answer, you need to know whether the retriever failed to find the right document, the re-ranker sorted it too low, the chunker split the relevant content across boundaries, or the LLM ignored correct context and hallucinated anyway.

Our evaluation framework measures four components independently. Retrieval quality uses recall@k and NDCG on a curated set of 2,000 query-document relevance judgments. Re-ranking quality measures the precision improvement between raw retrieval and re-ranked results. Generation faithfulness uses an NLI model to check whether each response claim is supported by the retrieved context. Answer relevance measures whether the response actually addresses the user's question.

We run automated evaluations on every pipeline change and weekly on a rolling sample of production queries. The automated metrics correlate at 0.78 with human judgments, which is good enough for regression detection but not for absolute quality assessment. Human evaluation on 300 queries per week catches the remaining quality issues that automated metrics miss.

The evaluation infrastructure connects to our broader experiment tracking and model monitoring systems, ensuring that pipeline changes are measured against consistent baselines. We've also integrated retrieval evaluation with our data quality checks to catch corpus-level issues before they affect retrieval.

Caching and Latency Optimization

Production RAG latency budgets are tight. Users expect sub-3-second responses, and the LLM generation alone consumes 1-2 seconds. That leaves roughly 800ms for the entire retrieval pipeline: query transformation, hybrid search, RRF merge, and re-ranking.

We cache at three levels. Embedding cache stores pre-computed query embeddings keyed on normalized query text, saving 20-50ms per query. Result cache stores full retrieval results with a 10-minute TTL, handling repeated and similar queries. Response cache stores complete generated answers for exact query matches, with invalidation triggered by corpus updates.

The re-ranker is the latency bottleneck in retrieval. We optimized it with quantized inference and dynamic batching, reducing re-ranking latency from 400ms to 180ms. For further optimization, we implemented speculative re-ranking: start LLM generation with the top-5 bi-encoder results while the cross-encoder re-ranks the full candidate set. If re-ranking changes the top-5, we interrupt and restart generation. This saves 200ms on 70% of queries where re-ranking doesn't change the top results.

The caching and latency strategies tie into general principles of model serving optimization and real-time feature serving at the infrastructure level.

Indexing Pipeline: Keeping the Corpus Fresh

A RAG system is only as good as its index. Stale documents, missing updates, and inconsistent embeddings degrade retrieval quality silently — the system doesn't error, it just returns increasingly irrelevant results.

Our indexing pipeline processes document updates in near-real-time using a change data capture stream. When a document is created or modified, the pipeline chunks it, computes embeddings, updates the vector index, and refreshes the BM25 index. The entire flow completes within 30 seconds of the source document changing.

Embedding model updates are more disruptive. When we retrain or swap the embedding model, all existing vectors become incompatible. We handle this with blue-green indexing: build the complete new index in parallel, validate it against the evaluation set, then atomically swap the active index. The old index stays available for rollback for 48 hours. This approach mirrors patterns from data versioning and canary deployment strategies.

Lessons from Production

After running RAG systems in production for two years, several lessons stand out. First, retrieval quality matters more than generation model quality. Upgrading from a mediocre retriever to an excellent one improved answer accuracy by 30%. Upgrading the LLM improved it by 8%. Second, your evaluation framework is the most important piece of infrastructure — without it, you're flying blind on every change.

Third, chunking is the foundation that everything else builds on. Bad chunks produce bad embeddings that produce bad retrieval that no amount of re-ranking or prompt engineering can fix. Invest in structure-aware chunking early. Fourth, users don't tell you when the system is wrong — they just stop using it. You need proactive quality monitoring, not just a feedback button.

Finally, start with the simplest pipeline that works and add complexity only when metrics justify it. Query transformation, hybrid retrieval, re-ranking, and guardrails each add latency and operational complexity. Measure the improvement from each component independently before committing to it in production. The CI/CD pipeline should validate each component's contribution on every deployment.