Kartik Aneja
Kartik Aneja

AI/ML Platform Engineer in Boston. I build production AI systems, streaming data platforms, and the MLOps glue between them.

Shipping a production RAG pipeline: the parts that matter.

A short write-up on the choices that tend to make or break retrieval systems once they move beyond a prototype: retrieval quality, caching, observability, and service boundaries.

Retrieval quality

You only get a good answer if the retrieval layer is predictable and well scoped.

Caching

A small cache can dramatically cut latency for repeated or similar requests.

Observability

Track latency, recall, and relevance before users tell you the system feels slow.

The main idea: production RAG is less about “adding embeddings” and more about designing the retrieval path so it stays debuggable, measurable, and affordable under load.
A concrete example: docubrain takes the observability point literally — every answer must cite the exact page it came from, enforced at the data-model level, or nothing is returned at all. Source →