How to Design Graders for AI Agent Evaluations
Reliable agent evaluation combines deterministic checks, model-based grading and human assessment, then tests whether each grader measures the intended behaviour consistently.
Match the grader to the claim
Use deterministic code for exact states, calculations, schemas, tool parameters, permissions and completion. Use rubric-based model grading when several outputs may be valid but quality is observable. Use people where consequence, novelty, context or value judgement makes automated scoring insufficient.
Write observable scoring anchors
Define what a pass, partial result and failure look like in behaviour. Distinguish not observed from not applicable. Avoid labels such as good reasoning unless the rubric identifies evidence that an assessor can inspect.
Give graders the right evidence
A final answer grader cannot judge tool misuse or an authority breach hidden in the trace. Provide the case, permitted evidence, relevant trace fields, tool outcome and business state needed for the criterion. Minimise unnecessary personal or confidential data.
Calibrate model-based graders
Compare the model grader with reviewed examples, difficult boundary cases and independent human ratings. Test sensitivity to answer order, irrelevant detail and stylistic polish. Use a separate grader model or blind configuration where practicable.
Measure agreement and investigate disagreement
Have more than one assessor independently score a sample. Use an agreement statistic suited to the scale and inspect the cases where ratings diverge. High agreement does not prove that the rubric is valid, but weak agreement makes the score difficult to trust.
Defend against grader gaming
Agents may exploit an incomplete rule, imitate a preferred phrase or manipulate a model judge. Combine outcome verification, trace review and adversarial cases. NIST distinguishes solution contamination and grader gaming as material risks for agent evaluations.
Version graders with the eval suite
A changed rubric or grader can alter scores even when the agent is unchanged. Version instructions, reference answers, thresholds and evaluator models. Re-score a stable calibration set before comparing results across time.
