Shipping a production RAG pipeline: the parts that matter.
A short write-up on the choices that tend to make or break retrieval systems once they move beyond a prototype: retrieval quality, caching, observability, and service boundaries.
Retrieval quality
You only get a good answer if the retrieval layer is predictable and well scoped.
Caching
A small cache can dramatically cut latency for repeated or similar requests.
Observability
Track latency, recall, and relevance before users tell you the system feels slow.
The main idea: production RAG is less about “adding embeddings” and more about designing the retrieval path so it stays debuggable, measurable, and affordable under load.