Guide

How to Evaluate AI Agent Memory and Multi-Step Follow-Through

Long-horizon evaluation tests whether an agent preserves relevant state, revises it when evidence changes and completes commitments across several interactions and events.

Memory is useful only when it serves the task

An agent may need to retain a customer's constraint, a pending approval or the result of an earlier tool call. It should not retain irrelevant or prohibited information. Evaluation begins by stating which facts must persist, for how long, in which scope and under whose authority. The expected behaviour includes forgetting, correction and deletion as well as recall.

Reveal important information gradually

Give the agent an initial task, then introduce a changed address, cancelled approval, delayed event or corrected policy later. Test whether it updates the relevant belief without overwriting unrelated facts. Concealed signals are useful when discovery is part of the workflow, such as noticing that two records refer to the same case or that a later event invalidates an earlier plan.

Test commitments rather than conversational recall

Remembering a fact in a response is different from following through. Record promised actions, owners, conditions and deadlines. Then test whether the agent resumes the right case, checks the latest state, completes the action or escalates when completion is impossible. An agent should not claim that a task is complete because it remembers intending to do it.

Separate memory failures from tool and workflow failures

A missed follow-up may arise because the state was not stored, the retrieval query was wrong, the event never arrived or the workflow had no resumption mechanism. Inspect traces and controlled state to locate the break. Rerun the case after changing one condition at a time where necessary.

Protect scope and permissions

Test that information from one customer, case, tenant or user session does not leak into another. Change roles and revoke access during a long-running case. The agent should respect the current permission even when an earlier state allowed the action. Sensitive memory requires explicit retention, deletion and review rules.

Measure eventual completion and recovery

Useful measures include correct state recall, successful correction, unresolved commitment rate, time to completion, duplicate action, inappropriate persistence and recovery after interruption. Repeat variable cases and preserve critical privacy or authority failures as separate gates rather than averaging them into a general memory score.

What to read next

References

  1. DeepEval, Agentic evaluation metrics
  2. Galileo, Evaluate documentation
  3. LangSmith, Evaluation types