Zero to AI engineer · Stop 11 of 14 lessons
Evaluating a RAG system without lying to yourself
What this lesson covers
The reason most RAG systems quietly get worse over time is that nobody’s measuring retrieval — only the final answer, which the model can compensate for long past the point that the retrieval is broken.
The outline
- The two-layer split. Retrieval quality (did we find the right chunks?) vs. answer quality (did we say the right thing?). Why you need both.
- The golden set. Fifty question-answer pairs you write by hand, with the source chunk labelled. Why fifty is enough and why more is often less.
- Retrieval metrics. Recall@k, MRR, and the one metric that actually correlates with users being happy.
- Answer metrics. LLM-as-judge, done right. The rubric shape and the inter-judge check that keeps you honest.
- The change that made retrieval worse. How to spot it in the numbers, why it’s the most common failure mode after month one.
- Continuous evaluation. Rebuilding the golden set from real traffic, without a labelling team.
Coming soon
In outline.