Most RAG demos hide the hard parts: data freshness, permission boundaries, embedding drift, and cold-start latency on GPU nodes. Production systems need an explicit checklist before traffic arrives.
Separate ingest from query paths. Index rebuilds should be versioned and switchable. Document ACLs must survive chunking—retrieval that ignores identity is a data leak waiting to happen.
On Kubernetes, pin GPU node pools, set request/limit discipline, and watch queue depth on inference services. Pair OpenTelemetry traces across the API, retriever, and model endpoint so you can tell a slow vector store from a slow model.
Finally, measure answer quality continuously. Offline eval sets catch regressions; online feedback catches real user failures.
Add evaluation fixtures that catch citation failures before customers do.
Rate-limit and budget token spend per tenant or product surface.
Keep PII out of prompts and indexes unless retention, access logging, and redaction are designed on purpose.
Key takeaways
- Treat retrieval indexes as versioned production data with rollback.
- Separate embedding jobs, API serving, and vector stores for blast-radius control.
- Measure groundedness and latency together—accuracy without SLOs is incomplete.
FAQ
Should the vector DB run inside the cluster?
It can, but managed stores often win on backups and ops. If in-cluster, define storage, restore drills, and resource isolation explicitly.
What breaks first in production RAG?
Stale indexes, embedding model drift, secret sprawl for model APIs, and unbounded context that blows latency and cost.