Elo Rating

An Elo rating is a relative skill score derived from head-to-head comparisons, borrowed from chess and applied to models. In systems like Chatbot Arena, human voters pick the better of two model responses, and Elo converts these pairwise wins and losses into a ranking that reflects overall preference.

Definition

Elo Rating

Elo scores rank competitors by outcomes against one another rather than by an absolute test. Each comparison updates the participants' ratings based on who won and how surprising the result was, so a model that consistently beats strong opponents rises. This suits subjective quality, where no fixed answer key exists.

Chatbot Arena popularized Elo for language models by crowdsourcing blind pairwise votes between anonymized model outputs. The approach captures human preference at scale, but reflects the voter population and prompt distribution, so it complements rather than replaces task-specific benchmarks and evals.

Frequently asked questions

What is Elo Rating?

An Elo rating is a relative skill score derived from head-to-head comparisons, borrowed from chess and applied to models. In systems like Chatbot Arena, human voters pick the better of two model responses, and Elo converts these pairwise wins and losses into a ranking that reflects overall preference.

Sources

Related terms

← Full AI & prompt engineering glossary

Put Elo Rating to work with Prompeteer, the Agentic Contextual AI Platform →