Guide

Continuous AI Agent Evaluation in Production

Production evaluation detects new cases, regressions and operating drift by connecting sampled traces, incidents, business outcomes and versioned offline tests.

Separate offline and online evaluation

Offline evals run controlled cases before release. Online evaluation samples real traces and outcomes after deployment. Online signals discover change; controlled offline tests determine whether a candidate fix improves the system without breaking supported behaviour.

Trace every material component

Record the agent and workflow version, input, retrieved sources, tools, permissions, handovers, approvals, outputs, external state and later correction. Use identifiers and access controls that support investigation without copying unnecessary sensitive content.

Sample by risk and outcome

Review a random sample for broad quality and targeted samples of overrides, abstentions, errors, high-cost traces, unusual routes, complaints and consequential actions. Averages alone may miss the part of the workflow that changed.

Use automated evaluators as signals

Run deterministic checks on permissions, schemas, tool outcomes and completion. Apply calibrated model graders to suitable quality criteria. Route low-confidence, high-severity and novel cases to people. Automated scores support triage; they do not remove accountable review.

Turn failures into versioned tests

Reconstruct incidents, near misses and corrected cases in an authorised test environment. Add them to the regression suite with the relevant source versions and acceptance rule. Test the fix and adjacent behaviour before release.

Set change and stop rules

Trigger full or targeted suites after model, prompt, retrieval, tool, permission, policy or workflow changes. Define conditions that restrict authority, revert a release or pause the route. Preserve manual continuity while the issue is investigated.

Review business performance over time

Track outcome, cycle time, rework, customer effort, human review, cost and downstream corrections. Compare like periods and case mixes. A technically stable agent may become unsuitable when the process, policy or customer population changes.

Continue through the agent evaluation series

References

  1. Microsoft, Agent evaluation checklist
  2. Microsoft, Agent evaluators
  3. NIST, TEVV-Athlon framework