5 · Production (MLOps + LLMOps) › Phase 6 › Lesson 1 of 5
Offline evals that catch regressions
What this lesson covers
Evals are the single most important thing in the production track. They’re also the thing most teams get subtly wrong — the case file eval-pipeline-that-lied is what happens when you don’t.
The outline
- The three parts. A dataset, a set of graders, a scoring aggregation. Every eval harness is a variation on these three.
- The golden set. How to build one in a day. Why 50–200 examples beat 1000 auto-generated ones.
- Graders that work. Regex, exact match, LLM-as-judge, rubric-based, pairwise. When to use each.
- LLM-as-judge, honestly. The prompt template, the inter-judge check, and the failure modes that quietly poison your metric.
- The CI gate. Wiring the harness into a pre-merge check, blocking on regressions, and the cost budget for the gate itself.
- When your green eval is lying. How to check whether your eval is still measuring what you think.
Coming soon
In outline.