Production RAG: hybrid retrieval, eval, and when to skip it
Most RAG tutorials stop at “embed chunks, cosine similarity, paste into the prompt.” That works in a notebook. In production it fails quietly: answers sound confident, citations are wrong, latency spikes on Monday mornings, and nobody can tell whether the last prompt change or the last reindex caused the regression.
This study is about the engineering layer after the demo: retrieval design, hybrid search, evaluation harnesses, and the decision to not use RAG at all.
What RAG is actually buying you
RAG is not “the model knows your docs.” It is bounded context injection: at answer time you fetch a small set of passages and ask the model to stay inside them. You trade:
- Pros: fresher knowledge without full fine-tunes, auditable sources, per-tenant isolation in the index.
- Cons: retrieval errors become answer errors; token budget caps how much you can inject; embedding and chunking choices are part of the product.
Chunking is a product decision
There is no universal chunk size. What you optimize for depends on how users ask questions.
| Strategy | Good when | Failure mode | |----------|-----------|--------------| | Fixed token windows | Uniform docs, FAQ-like queries | Splits tables and procedures mid-step | | Paragraph / heading-aware | Policies, runbooks, markdown KB | Short sections with no local context | | Parent–child (small retrieve, large read) | Long PDFs, legal or spec text | More storage and join logic at read time | | Metadata-tagged chunks (product, region, version) | Multi-catalog support | Bad metadata → silent wrong scope |
Overlap helps continuity; it also multiplies index size and can duplicate contradictory paragraphs from different doc versions. Versioning source documents and reindex idempotently (delete old vectors for document_id, then insert) beats append-only chaos.
Practical rule: log the chunk ids that entered the prompt. When support says “the bot lied about warranty,” you replay that trace in minutes.
Embeddings drift without drama
Same text, new embedding model, new vector space — old and new vectors must not coexist in one index unless you dual-write and cut over explicitly. Production checklist:
1. Pin embedding model id in config (text-embedding-3-small, etc.).
2. Reindex job is observable: documents/hour, failures, dimension mismatch guards.
3. Similarity thresholds are per corpus, not copied from a blog post. A threshold of 0.55 means nothing without your score distribution.
When quality drops after a model vendor update, assume retrieval moved before you rewrite the system prompt.
Hybrid search: when vectors alone are not enough
Pure dense retrieval misses exact tokens: SKUs, error codes, internal project names, “section 4.2.1”. Hybrid patterns that hold up:
- Dense + BM25 (or Postgres
tsvector) with reciprocal rank fusion or learned weights. - Pre-filter on metadata (
organization_id,product_line,doc_status=published) before similarity — multi-tenant RAG that skips this will leak context across tenants in subtle ways. - Query rewriting (HyDE, sub-queries) only after you measure baseline recall; extra LLM calls add latency and new failure modes.
Grounding and abstention
A retriever that returns nothing useful should not trigger creative completion. Patterns that survive audit:
- Explicit “no relevant sources” branch: offer human handoff or narrow clarification, not a guess.
- Require the model to cite chunk ids or source titles in internal tools; reject answers that cite ids not in the context block (programmatic check where feasible).
- Cap injected context size; truncate with structure (“Source 1: …”) so the model can refer consistently.
Evaluation before and after every change
If you only evaluate by reading ten chats in Slack, you will ship regressions. Minimum viable eval loop:
1. Golden set: real user questions (redacted) + expected supporting doc ids or answer rubric. 2. Retrieval metrics: recall@k, MRR on labeled q→doc pairs; separate from generation quality. 3. Generation metrics: LLM-as-judge with human spot-check, or human-only for high-risk domains. 4. CI gate: small golden set on PRs that touch chunking, embeddings, or retrieval SQL.
Regression in recall@5 is a blocker; tweaking the system prompt to “sound nicer” while recall dropped is how teams lose trust.
When not to use RAG
RAG is the wrong tool when:
- The knowledge is small and stable — fit it in the system prompt or a compiled tool schema.
- You need deterministic computation (pricing, eligibility) — use code or a rules engine; let the LLM explain results.
- Latency SLO is tight and corpus is huge — consider precomputed summaries, routing to specialized sub-agents, or fine-tuned small models on a frozen snapshot.
- Security classification forbids mixing passages in one prompt — partition indexes and route queries first.
Operating the stack
What senior teams instrument:
- p95 retrieval latency and embedding batch duration
- Empty-hit rate and below-threshold rate per org
- Token count of assembled context vs model limit
- Traces tying
request_id→ chunk ids → model completion id
Closing
Production RAG is search engineering plus policy plus observability, with an LLM at the end. Chunking, hybrid retrieval, tenant-safe filters, abstention, and a golden eval set are what separate a demo from something legal and ops will run.
If you already run a product-specific stack (voice tools, inbox, Kafka-backed workers), keep one retrieval spine and hang channels off it — duplicate retrievers per surface is how divergence starts.