Interview Q&A · 5 · Production (MLOps + LLMOps)
What's the difference between an eval and a benchmark, and why would a team run both?
Reveal the answer
A benchmark is a fixed public dataset scored by a public rubric. It's
useful for comparing models against each other on the *same* task.
MMLU, HumanEval, GSM8K, MT-Bench — all benchmarks. Numbers are
comparable across teams; they're also aggressively Goodharted, so a
high benchmark score is not proof a model is good for *your* use case.
An eval is a private, task-specific measurement of "does the feature I'm
actually shipping still work when I change something?" It's a golden set
of your inputs, a set of graders you wrote (or an LLM-as-judge with a
rubric you wrote), and a threshold you set. A green eval is a green
light for a specific ship; it says nothing about which model is best in
general.
A team runs both because they answer different questions. Benchmarks:
"should we consider swapping to model X?" Evals: "did this prompt change
we just made regress the answer for our users?" The bug most teams have
is running only benchmarks (and shipping regressions) or only evals
(and missing that a new model would solve the problem entirely).
Common variants
- How do you build a golden set for a task with no ground truth?
- When is LLM-as-judge better than exact-match, and when is it a trap?
- How would you detect that your own eval has drifted from what users care about?
Track this card
Stored in this browser only — no signup, no sync. Clearing site data clears it.