Elo Rating
An Elo rating is a relative skill score derived from head-to-head comparisons, borrowed from chess and applied to models. In systems like Chatbot Arena, human voters pick the better of two model responses, and Elo converts these pairwise wins and losses into a ranking that reflects overall preference.
Definition
Elo RatingElo scores rank competitors by outcomes against one another rather than by an absolute test. Each comparison updates the participants' ratings based on who won and how surprising the result was, so a model that consistently beats strong opponents rises. This suits subjective quality, where no fixed answer key exists.
Chatbot Arena popularized Elo for language models by crowdsourcing blind pairwise votes between anonymized model outputs. The approach captures human preference at scale, but reflects the voter population and prompt distribution, so it complements rather than replaces task-specific benchmarks and evals.
Frequently asked questions
What is Elo Rating?
An Elo rating is a relative skill score derived from head-to-head comparisons, borrowed from chess and applied to models. In systems like Chatbot Arena, human voters pick the better of two model responses, and Elo converts these pairwise wins and losses into a ranking that reflects overall preference.
Sources
Related terms
← Full AI & prompt engineering glossary
Put Elo Rating to work with Prompeteer, the Agentic Contextual AI Platform →