Guide

How to Evaluate AI Agent Tool Use and Trajectories

Tool and trajectory evaluation tests whether an agent selected the right action, supplied valid inputs, respected permissions, used the result correctly and verified completion.

Evaluate the complete action chain

Inspect the reason a tool was selected, its inputs, permission context, returned result, interpretation, next action and external state. The chain should remain traceable from user intent to verified outcome.

Test selection and arguments separately

An agent can choose the correct tool with the wrong account, date, currency or limit. Score necessity, selection and parameter accuracy independently. Test omitted required calls and redundant calls that create cost or duplicate effects.

Distinguish acknowledgement from completion

Many systems accept a request before the business action finishes. The agent should retain a pending state, check the authoritative record and avoid telling the user that work is complete without evidence.

Introduce failures and partial success

Test timeouts, rate limits, expired credentials, unavailable systems, incomplete responses and one successful step followed by a failed step. Verify retry limits, idempotency, compensation and unresolved-case ownership.

Test permissions and confirmation rules

Attempt actions beyond the agent's role, just above thresholds and through split transactions. Test missing approvers and misleading requests to bypass controls. The system should decline or escalate without executing the prohibited action.

Allow different valid paths

Several trajectories may achieve the same outcome safely. Score required constraints, efficiency and unnecessary risk instead of demanding one exact sequence. Investigate unusually short paths that may have skipped evidence or controls.

Turn production traces into regression cases

Sample corrected, overridden, abandoned and unusually expensive traces. Reconstruct the relevant environment and add them to a versioned suite. Re-run the suite after any change to tools, prompts, models, policies or integrations.

Continue through the agent evaluation series

References

  1. Microsoft, Agent evaluators
  2. Anthropic, Demystifying evals for AI agents