Evaluation set
An evaluation set is a collection of real inputs, each paired with what a good output looks like, used to score an AI system before it's released. It is to an AI product what a test suite is to ordinary software.
What it means
AI systems don't give the same output every time, and a change that fixes one case can quietly break others. An evaluation set catches that: every change is scored against the same cases before it ships, so a release that makes things worse never reaches users.
A good evaluation set is built from real inputs, not invented ones, and includes the hard cases: the ambiguous request, the input designed to mislead, the question the system should refuse. It grows every time a new failure is found.
How we use it
Our method In our product builds, the evaluation set is designed before the build starts, alongside the success metric, so the build has something to be measured against from the first week. It's one of the decisions in how to scope an AI product.
Related terms
Published 28 September 2026. All terms