Benchmark

A benchmark is a standardized dataset and scoring method used to measure and compare model capabilities on a defined task, such as reasoning, coding, or question answering. Shared benchmarks make progress comparable across models, though results can be inflated by data contamination or narrow test design.

Definition

Benchmark

Benchmarks provide a common yardstick: everyone runs the same inputs and scores them the same way, so models can be ranked on a capability. Well-known examples include MMLU for broad knowledge and GSM8K for grade-school math reasoning. They are essential for tracking field-wide progress and for choosing models.

Benchmarks have limitations. A model may have seen test data during training (contamination), a benchmark may not reflect real-world tasks, and optimizing to a benchmark can overstate general ability. For this reason teams complement public benchmarks with their own task-specific evals.

Frequently asked questions

What is Benchmark?

A benchmark is a standardized dataset and scoring method used to measure and compare model capabilities on a defined task, such as reasoning, coding, or question answering. Shared benchmarks make progress comparable across models, though results can be inflated by data contamination or narrow test design.

Sources

Related terms

← Full AI & prompt engineering glossary

Put Benchmark to work with Prompeteer, the Agentic Contextual AI Platform →