From Production Failure to Reproducible AI Agent Regression Test
A production incident becomes useful evaluation evidence only after the failure has been verified, reduced to a reproducible case and protected by a permanent release test.
Preserve the incident before changing the system
Capture the request, relevant conversation, retrieval results, tool calls, model and prompt version, external state, permissions, event timing and final outcome. Preserve sensitive information carefully and record missing evidence. Changing the prompt immediately may remove the symptom without explaining the failure or proving that the fix works.
Establish what should have happened
Review the applicable policy, workflow and authority at the time of the incident. Distinguish a system failure from an ambiguous requirement or incorrect operational data. Record the required, permitted and prohibited behaviours. A historical human action is useful context, but it is not automatically the correct reference.
Reproduce the failure in a controlled environment
Recreate the necessary records, tool responses and events in a sandbox or mock environment. Confirm that the candidate version fails for the same material reason. If the incident cannot be reproduced, retain it as an investigation case and avoid claiming that the cause has been proven.
Reduce the case without removing the cause
Remove irrelevant details one at a time until the smallest reliable reproducer remains. Change suspected conditions selectively to test the diagnosis. The process may show that stale evidence, a permission mismatch, a retry, an unclear instruction or an interaction between components caused the outcome.
Test the correction and neighbouring behaviour
Rerun the reproducer after remediation and test adjacent cases that could be affected. A stricter refusal rule may stop the incident while making ordinary work unusable. Compare the fix with the approved baseline using the same environment, grader versions and repetitions.
Keep the case as a permanent release gate
Add the verified case to the regression suite with its starting state, expected behaviour, severity and ownership. Link it to the incident without exposing unnecessary production data. Run it whenever the model, prompt, tool, policy, permission or related workflow changes. Periodically confirm that the fixture still represents the current business rule.
Example: a duplicate payment after an uncertain tool response
A finance agent submits a permitted payment request and the endpoint times out. The agent retries, creating a second instruction, then tells the operator that the supplier has been paid. The incident record should preserve both request identifiers, tool responses, ledger states, permissions and timing. In a sandbox, the evaluator seeds the same invoice and makes the first call succeed while returning an uncertain response. The expected behaviour is to check authoritative state and escalate if it cannot be established. The regression test fails on any duplicate instruction or false settlement statement.
The permanent regression asset
Marketways converts the verified incident into a versioned fixture containing the minimum starting records, event schedule, mock responses, expected behaviour, prohibited mutations and severity. The fixture links to the incident without copying unnecessary production data. It runs when the model, prompt, payment tool, retry logic, policy or permission changes. Results are compared with neighbouring ordinary cases so the correction does not prevent legitimate processing. Ownership and a review date keep the test aligned with the current finance control.
What to read next
- Continuous agent evaluation in production: find and sample new operating evidence
- Controlled environment testing: reproduce the incident safely
- Build an evaluation test set: integrate verified failures without distorting coverage
