How to Build an AI Agent Evaluation Test Set
A strong agent evaluation set represents the real workflow, material failure modes, ordinary cases, rare consequences and plausible false alarms.
Start with the operating population
Define the users, requests, cases, systems, languages, channels and time period the agent will encounter. Use process data, reviewed cases, incident records, interviews and observation. A convenient list of prompts is not automatically representative of the work.
Create a failure-mode inventory
For each workflow stage, ask how evidence, routing, reasoning, tools, permissions, memory, handovers and completion could fail. Include interactions between failures, such as stale policy combined with an unavailable approver.
Balance ordinary, difficult and control cases
Include routine work to measure usefulness, genuine exceptions to test detection and plausible false alarms to test restraint. Add ambiguous cases with several defensible responses. Record which behaviour is required, permitted or prohibited.
Preserve the evidence available at each step
Version documents, records, tool responses and policy. The agent should be judged on what it could know at the time, not information added after the outcome. This is essential when testing multi-turn discovery and delayed events.
Separate development and held-out cases
Developers need cases for iteration, while release decisions need cases that have not shaped the system. Limit access to sensitive test sets where leakage would make the score meaningless. Add new holdouts as the system and workflow change.
Specify repetitions and comparison conditions
Identify which cases require repeated trials and hold the model, tools, data and environment constant when comparing versions. Record every component version so an improvement or regression can be investigated.
Document coverage and limits
Map every test to workflow route, failure mode, consequence and population. State what remains untested. A test set gives evidence about its coverage; it does not prove reliability for every future case.
