How to Evaluate AI Agent Tool Use and Trajectories
Tool and trajectory evaluation tests whether an agent selected the right action, supplied valid inputs, respected permissions, used the result correctly and verified completion.
Evaluate the complete action chain
Inspect the reason a tool was selected, its inputs, permission context, returned result, interpretation, next action and external state. The chain should remain traceable from user intent to verified outcome.
Test selection and arguments separately
An agent can choose the correct tool with the wrong account, date, currency or limit. Score necessity, selection and parameter accuracy independently. Test omitted required calls and redundant calls that create cost or duplicate effects.
Distinguish acknowledgement from completion
Many systems accept a request before the business action finishes. The agent should retain a pending state, check the authoritative record and avoid telling the user that work is complete without evidence.
Introduce failures and partial success
Test timeouts, rate limits, expired credentials, unavailable systems, incomplete responses and one successful step followed by a failed step. Verify retry limits, idempotency, compensation and unresolved-case ownership.
Test permissions and confirmation rules
Attempt actions beyond the agent's role, just above thresholds and through split transactions. Test missing approvers and misleading requests to bypass controls. The system should decline or escalate without executing the prohibited action.
Allow different valid paths
Several trajectories may achieve the same outcome safely. Score required constraints, efficiency and unnecessary risk instead of demanding one exact sequence. Investigate unusually short paths that may have skipped evidence or controls.
Turn production traces into regression cases
Sample corrected, overridden, abandoned and unusually expensive traces. Reconstruct the relevant environment and add them to a versioned suite. Re-run the suite after any change to tools, prompts, models, policies or integrations.
