Evaluation (Evals)

Evaluation, or "evals," is the systematic measurement of a model or AI application's quality against defined criteria and test cases. Evals quantify accuracy, safety, format compliance, and task success, turning subjective impressions into repeatable metrics that guide prompt, model, and system changes.

Definition

Evaluation (Evals)

Because model behavior is probabilistic and sensitive to prompt changes, teams need objective measurement to know whether a change helps or hurts. Evals define representative inputs and scoring — exact match, rubric grading, human review, or an LLM acting as a judge — and run them repeatedly as prompts, models, or retrieval change.

Good evaluation covers correctness, robustness, safety, and format, and is run as a regression suite so improvements are verified rather than assumed. Benchmarks are shared, standardized evals used to compare models across the field.

Frequently asked questions

What is Evaluation (Evals)?

Evaluation, or "evals," is the systematic measurement of a model or AI application's quality against defined criteria and test cases. Evals quantify accuracy, safety, format compliance, and task success, turning subjective impressions into repeatable metrics that guide prompt, model, and system changes.

Sources

Related terms

← Full AI & prompt engineering glossary

Put Evaluation (Evals) to work with Prompeteer, the Agentic Contextual AI Platform →