AI Vocabulary
Benchmarks: do published scores say anything about your use?
A benchmark is a standardised test comparing models on the same set of questions. Published scores measure general capability, on short isolated tasks, in conditions unlike those of a matter. They are useful to laboratories and of little help in choosing a tool.
What they actually measure
A series of questions with verifiable answers, asked without context, of which correct answers are counted. Existing legal tests often draw on professional examinations: closed questions, a single answer, a complete set of facts supplied.
Those are conditions in which a model excels, and they have almost nothing in common with real work. An examination question contains everything needed to answer it; a matter never has that property.
Four reasons for caution
Contamination: a published test ends up in the training corpora of later models, which then pass it without demonstrating anything.
Format mismatch: scores cover short tasks, while difficulties appear on long ones, where definitions and cross-references must be held.
Absence of context: nothing measures the ability to work on a matter of several hundred documents with earlier decisions.
Choice of test: a provider publishes the benchmarks on which it does well, which is legitimate and makes comparison between announcements inconclusive.
A case that shows the gap
A model passing a professional examination with a high score is impressive, and the result is real. But look at the conditions: the question contains every relevant fact, it calls for a single answer, and nobody asks the candidate to cite sources.
Now take the same question on a real matter. The facts are spread across 80 documents, some of which contradict one another. Several answers are defensible depending on the strategy adopted. And every assertion will have to be tied to a verifiable source.
The examination score says nothing about the second exercise, which is the one you are buying.
What it does not solve
A high score says nothing about what concerns you: the reliability of cited references, the coherence of a long document, the behaviour when the system cannot find what it is looking for.
And a low score does not disqualify. A model that is middling on tests, properly framed and supplied with your matter, will produce better work than an excellent model interrogated in a vacuum.
What replaces a benchmark
A test run at your firm, on a matter you know and a question to which you already have the answer. It takes longer to organise than reading scores, and it is the only measure that bears on your use.
Three conditions make it probative: that the matter be real and recent, that the question have been dealt with already so the result can be judged at once, and that it be difficult — an easy question proves nothing.
Why it matters to a lawyer
Because these scores are the most frequent argument in meetings, and the hardest to contradict without appearing hostile to measurement. Yet contesting them is the rigorous position here: a measurement taken in conditions foreign to yours does not predict your result.
The answer to offer is simple and hard to refuse: let us test it on one of our matters.