Machine Learning Metrics vs AI Agent Evals: Accuracy Is Only the Beginning
Precision, recall and calibration remain useful, but an agent also needs evaluation of trajectories, tools, authority, handovers and verified business outcomes.
Classical metrics answer a defined prediction question
A classifier can be assessed by comparing predicted and actual classes. Accuracy reports the proportion correct. Precision describes how often a positive prediction is correct, recall how many actual positives are found, and a confusion matrix shows which classes are confused. Calibration asks whether stated probabilities correspond to observed frequencies.
Agent tasks have several linked opportunities for error
An agent may correctly classify a case and still retrieve the wrong policy, choose the wrong tool, pass an invalid parameter, exceed its authority or leave the work incomplete. These are workflow and control errors rather than prediction errors. They require traces, system state and business evidence.
Retain class-level analysis where it matters
Agent evals should still use false-positive and false-negative analysis for routing, detection, escalation and abstention. Aggregate accuracy can hide a severe class such as a missed safety complaint or fraudulent payment. Report results by case type, consequence and material subgroup.
Add process and outcome measures
Measure tool selection, argument accuracy, handover quality, duplicated actions, verified completion, latency, cost and human correction. Evaluate whether the result was achieved through an acceptable path. Several trajectories may be valid, so the goal is not to force one script but to detect unsafe or wasteful behaviour.
Use repeated trials for variable systems
Run important cases more than once and estimate the probability of each material outcome. Report denominators and uncertainty. A system that succeeds nine times and breaches authority once is different from a deterministic system with ninety per cent classification accuracy.
Connect metrics to failure costs
A missed urgent case, unnecessary escalation and slow completion have different costs. Set thresholds using consequence, prevalence, review capacity and the value of intervention. Keep critical controls as separate gates instead of hiding them inside an average.
