Evaluation & metrics
A shared, public set of tasks with a fixed way of scoring, so that different models can be compared on the same terms.
Formal
A published collection of tasks with known correct answers and a fixed scoring method, used across labs to rank models; well-known ones test broad school knowledge or the fixing of real faults in code.
In plain English
Like a standard fitness course that every athlete runs, with the same hurdles and the same stopwatch, so their times can be put side by side.
In practice
An IT architect in a ministry compares five large language models by their published benchmark scores, picks two, and then tests both on 300 real questions from the ministry's own staff before choosing.
Why it matters
Benchmarks drive how models are sold and ranked, but a model can be tuned to the questions, or have seen them in its training data, so a high score does not prove it will do well on your work.