aiengineering.guideaiengineering.guide

Case File · an offline LLM evaluation harness

The eval pipeline that lied for a month

Published

A team shipped model changes against a green eval suite for a month while quality quietly dropped. Here's the decomposition — and the wrong turn that hid it.

The constraint was simple: a small team wanted to change prompts and swap models without shipping regressions. They built an offline eval harness — a fixed set of inputs, a set of graders, and a CI gate that blocked a deploy if the aggregate score dropped. For a month, every change shipped green. Users, meanwhile, filed a steady trickle of complaints that answers had gotten worse.

The design

The harness had three parts. A dataset of ~300 representative inputs, captured from real traffic and frozen. A set of graders — mostly an LLM-as-judge scoring each answer 1–5 against a rubric, plus a few exact-match checks for structured fields. And a gate: the mean judge score had to stay within 0.1 of the last release or CI failed. It ran on every pull request in about four minutes.

The wrong turn

What actually fixed it

Three changes, in order of impact. First, the aggregate moved from a mean to a per-slice minimum: the dataset was tagged by difficulty and category, and the gate required no slice to regress, not just the average. The hard-input regression surfaced immediately. Second, the judge was pinned to a fixed model and version independent of the model under test, at temperature 0, and its own agreement with a small human-labeled set was tracked as a separate metric — so a drifting judge was now visible. Third, a canary of real traffic was scored in production and compared weekly against the offline number, because the frozen dataset had aged and no longer matched what users actually asked.

What to take from it

An eval suite is a measurement instrument, and an instrument that is never checked against ground truth will happily report whatever is cheapest to report. The failure here was not the LLM-as-judge — it was using the judge with no calibration, a single averaged number, and a dataset nobody re-validated. The green checkmark felt like safety; it was the absence of a signal. If your eval gate has never once caught a regression you also confirmed by hand, you do not yet know whether it works.

Verified · Aug 2026

Enter to go · Esc to close