aiengineering.guideaiengineering.guide

5 · Production (MLOps + LLMOps) › Phase 6 › Lesson 1 of 5

Offline evals that catch regressions

What this lesson covers

Evals are the single most important thing in the production track. They’re also the thing most teams get subtly wrong — the case file eval-pipeline-that-lied is what happens when you don’t.

The outline

  1. The three parts. A dataset, a set of graders, a scoring aggregation. Every eval harness is a variation on these three.
  2. The golden set. How to build one in a day. Why 50–200 examples beat 1000 auto-generated ones.
  3. Graders that work. Regex, exact match, LLM-as-judge, rubric-based, pairwise. When to use each.
  4. LLM-as-judge, honestly. The prompt template, the inter-judge check, and the failure modes that quietly poison your metric.
  5. The CI gate. Wiring the harness into a pre-merge check, blocking on regressions, and the cost budget for the gate itself.
  6. When your green eval is lying. How to check whether your eval is still measuring what you think.

Coming soon

In outline.

Outline

Enter to go · Esc to close