← Essays
Open question

A benchmark score is the start of a question, not the end

Performance numbers become meaningful only after we recover the choices hidden inside them.

A benchmark score looks like an ending. It appears at the far right of a table, often bolded, ready to settle a comparison.

I find it more useful to treat the score as the beginning of an investigation. The number is not false; it is compressed. The important work is recovering the decisions that produced it.

What population does the score describe?

An average can combine examples that differ in difficulty, source, language, format, or relevance to the intended use. Two systems with the same overall score may fail on entirely different parts of the distribution.

Before reading the ranking, I want to know:

  • how the examples were sampled;
  • whether categories are balanced or weighted by frequency;
  • which cases were excluded during cleaning; and
  • whether the tested population resembles the one behind the paper’s broader claim.

The unit of analysis matters too. An average over questions, users, documents, or repeated runs tells a different story.

What behavior became “correct”?

Metrics do not simply observe quality. They define a rule for recognizing it.

Exact match rewards one kind of behavior. Preference judgments reward another. A model-based grader introduces the grader’s own sensitivities. Even carefully designed human evaluation depends on instructions, examples, and the context raters receive.

So I try to rewrite every metric in behavioral language: what must a system do for this evaluation to count it as successful?

That translation often reveals a gap between the measured task and the conclusion attached to it.

How stable is the comparison?

A small score difference can feel decisive when the table hides uncertainty. I look for variation across random seeds, prompts, judges, subsets, and reasonable evaluation choices.

The useful question is not only “Is A higher than B?” It is also:

Across how many plausible versions of this evaluation would I still make the same decision?

This does not mean every paper needs an exhaustive robustness study. It means the strength of the conclusion should track the stability of the evidence.

What did improvement cost?

A better score may require more inference-time compute, extra tools, longer outputs, private data, or human filtering. Those costs can be acceptable. They simply belong in the comparison.

I want the result stated as a tradeoff, not a floating number: performance under a particular budget, latency, data regime, and level of supervision.

The question I keep

When I see a leaderboard, I now ask: what decision is this score supposed to support?

If the decision is choosing a system for a real setting, the most useful evaluation may not be the one with the broadest coverage. It may be the one whose population, failure costs, and measurement process most closely match that setting.

The score still matters. But the path from score to decision is the part I want to understand.

Correspondence

A counterexample is often more useful than agreement.

If this essay brings another paper, question, or objection to mind, I would be glad to hear it.

Write to me