Glossary · 6 · Quality, risk and governance
Evaluation (evals)
Also known as: Evals, LLM evaluation, Benchmark
Evaluation — “evals” for short — is the systematic measurement of how well an AI system performs a task, using a fixed set of test inputs with expected outputs or scoring criteria, scored by rules, by people or by another model, and repeated whenever the model, prompt or content changes.
- Expert
- Technical writers
- Technical project managers
- Developers
In one sentence
AI evaluation (evals) explained: measuring AI quality with test sets and scoring — the expert practice that makes AI projects controllable.
Example
Before launch, a team runs 200 real customer questions through its documentation assistant and checks whether each answer is correct, complete, cited and safe.
Why it matters on your learning path
- Technical writers: Writers are ideal authors of eval sets: they know the real questions, the right answers and the edge cases.
- Technical project managers: No evals, no release. Define quality thresholds as acceptance criteria and track them over time.
- Developers: Automate evals in CI; combine exact checks, model-graded scoring and periodic human review.
Benchmarks vs. your own evals
Public benchmarks compare models on general tasks. Only evals on your content and questions tell you whether a system works for your users — see model selection.