aiengineering.guideaiengineering.guide

Interview Q&A · 5 · Production (MLOps + LLMOps)

What's the difference between an eval and a benchmark, and why would a team run both?

mediumevalsproductionqualityasked at AnthropicCohereScale AI· 2026source: Case File · The eval pipeline that lied for a month

Reveal the answer
A benchmark is a fixed public dataset scored by a public rubric. It's useful for comparing models against each other on the *same* task. MMLU, HumanEval, GSM8K, MT-Bench — all benchmarks. Numbers are comparable across teams; they're also aggressively Goodharted, so a high benchmark score is not proof a model is good for *your* use case. An eval is a private, task-specific measurement of "does the feature I'm actually shipping still work when I change something?" It's a golden set of your inputs, a set of graders you wrote (or an LLM-as-judge with a rubric you wrote), and a threshold you set. A green eval is a green light for a specific ship; it says nothing about which model is best in general. A team runs both because they answer different questions. Benchmarks: "should we consider swapping to model X?" Evals: "did this prompt change we just made regress the answer for our users?" The bug most teams have is running only benchmarks (and shipping regressions) or only evals (and missing that a new model would solve the problem entirely).

Common variants

  • How do you build a golden set for a task with no ground truth?
  • When is LLM-as-judge better than exact-match, and when is it a trap?
  • How would you detect that your own eval has drifted from what users care about?

Track this card

Stored in this browser only — no signup, no sync. Clearing site data clears it.

Verified · Sept 2026

Enter to go · Esc to close