Statistical Design for AI Agent Evaluations: Repeated Trials and Reliable Conclusions
Agent evaluation needs sampling, repeated trials, uncertainty and assessor reliability so a score supports a defensible release decision.
Define the unit of analysis
Decide whether the observation is a user turn, complete case, workflow trajectory, customer, asset or repeated trial. Measurements from the same case or user are related and should not be treated as independent new evidence.
Sample the intended operating population
Represent common routes, important subgroups and rare but consequential conditions. Oversample a small high-risk group when necessary, then keep the design visible and avoid presenting the synthetic mix as the production rate.
Repeat variable cases
Run important cases several times to estimate stability. Preserve identical conditions where comparison is intended and record model sampling settings, tool state and external changes. Report pass proportions with denominators and intervals.
Compare versions on the same cases
Paired comparison reduces noise because each version faces the same evidence and constraints. Record model, prompt, retrieval, data, tool and workflow versions. Examine cases that improve and regress, not only the average difference.
Measure evaluator reliability
Use overlapping human ratings to estimate agreement. Test model graders against reviewed calibration cases and repeat grading where variability matters. Disagreement may reveal weak anchors, missing evidence or a genuinely ambiguous task.
Separate detection from intervention
Report missed issues and unnecessary interventions together. A system can appear highly sensitive by escalating everything. The preferred threshold depends on severity, base rate, review capacity, delay and the cost of acting or failing to act.
Avoid false precision
A test set may support a release boundary without estimating future production savings or failure prevalence. State coverage, exclusions and uncertainty. Use production evidence to update the conclusion rather than treating the laboratory score as permanent.
