Skip to content

GlossaryFloor 1 · The Modela solid block on its own: the prediction machineFloor 1 · The Model

benchmark

No. 061 · v2026-08FR: benchmark

A benchmark is a standardised test that scores models on a common set of questions, in order to compare them with one another. Like a national examination: it ranks the candidates, it does not say which one will do the job in your team.

What it is not

A benchmark is not an eval of your system. It measures a bare model on generic tasks, whereas what you put into service is a model plus a harness plus your documents plus your instructions: rankings do not travel. Nor is it a guarantee of quality: a high score sits perfectly well alongside systematic failures on your domain, and the gap between two neighbouring models in a ranking is often smaller than the measurement noise.

In depth

Contamination

The main problem with these tests is contamination. They are public, so they end up in training corpora, and a model can then restore answers it has seen rather than solve the problem. Benchmark authors respond by publishing renewed versions or keeping private sets, but the suspicion remains structural: the older and the more cited a test is, the less informative its score. A good reflex is to look at the date of a test as much as at the score obtained.

Selection bias

To this is added a selection bias in communication. A vendor picks the tests where its model stands out, and the conditions of measurement vary: number of attempts, format of the instructions, permission to reason at length. Two figures presented side by side may have been obtained under incomparable conditions. Rankings by human vote correct part of the problem but introduce another, since they measure preference, which rewards presentation as much as accuracy.

The practical conclusion

The practical conclusion is firm, and it is the one this lexicon defends everywhere: benchmarks serve to draw up a shortlist, never to decide. The choice is made on a set of tests of your own, built from your real cases, including those that have failed in the past, and replayed at every change of version. Thirty or so well-chosen cases say more about your use than any public ranking, and they keep their value when models change.

Relations where the neighbours live

Check 3 questions · click your answer

Level 1 · Recognise

A model comes first on the main benchmarks. Is it the right choice for your use?

Level 2 · Distinguish

What is meant by contamination of a benchmark?

Level 2 · Distinguish

What distinguishes a benchmark from an eval in the sense this site gives the word?

Who works with this 2 roles

The roles for which this term is part of the ordinary work.

No. 061 · v2026-08 · first written in · editorial responsibility Anthony Capirchio

Lexigraph, "Benchmark", v2026-08, https://www.lexigraph.org/en/benchmark/, CC BY 4.0.

Report

What goes with your message

Entry · Benchmark
No. 061 · v2026-08 · /en/benchmark

What is this about
0 / 600

It is used to reply to you, and for nothing else. What is recorded