aiengineering.guideaiengineering.guide

Zero to AI engineer · Stop 11 of 14 lessons

Evaluating a RAG system without lying to yourself

What this lesson covers

The reason most RAG systems quietly get worse over time is that nobody’s measuring retrieval — only the final answer, which the model can compensate for long past the point that the retrieval is broken.

The outline

  1. The two-layer split. Retrieval quality (did we find the right chunks?) vs. answer quality (did we say the right thing?). Why you need both.
  2. The golden set. Fifty question-answer pairs you write by hand, with the source chunk labelled. Why fifty is enough and why more is often less.
  3. Retrieval metrics. Recall@k, MRR, and the one metric that actually correlates with users being happy.
  4. Answer metrics. LLM-as-judge, done right. The rubric shape and the inter-judge check that keeps you honest.
  5. The change that made retrieval worse. How to spot it in the numbers, why it’s the most common failure mode after month one.
  6. Continuous evaluation. Rebuilding the golden set from real traffic, without a labelling team.

Coming soon

In outline.

Outline

Enter to go · Esc to close