Industry application

Worked Example: Evaluating a Multi-Agent Case with Persistent State

This illustrative example tests a multi-agent workflow in which evidence arrives over time, ownership changes and the system must complete a case without losing constraints or exceeding authority.

The proposed operating role

Assume an intake agent receives a request, an evidence agent gathers records, a specialist assesses the case and a coordinator completes or escalates it. The process can pause while waiting for a document or approval. The release decision concerns the complete workflow, not the local quality of any one agent. The multi-agent example is an illustrative application rather than a completed engagement.

The difficult conditions

The evaluation includes an incorrect early record that is later corrected, a delayed approval, conflicting specialist conclusions and a destination agent that becomes unavailable. One case revokes a user's permission while work is pending. Another introduces a concealed relationship between two records that the evidence agent should discover.

State and handoff evidence

The sandbox records case ownership, facts, provenance, pending commitments, permissions, handoff payloads and completed actions. Each handoff must state the task, relevant evidence, uncertainty and next required action. Shared state is separated from information that only an authorised specialist may access.

What the evaluation checks

The evaluator examines allocation, handoff completeness, correction of outdated facts, conflict detection, resumption after delay, permission enforcement and verified completion. It checks that no agent treats a summary as stronger evidence than the underlying record and that responsibility for the final outcome does not disappear between specialists.

How the finding changes the design

A failure may require a stronger handoff contract, explicit case owner, event-triggered resumption, provenance field or human reconciliation point. The corrected workflow is rerun against the original case and neighbouring cases. The recommendation states which coordination patterns are supported and which require further controls or a simpler single-agent design.

The principal practical limit

A multi-agent test can become complex enough to obscure the business question. The evaluation should add specialists only where the operating design genuinely requires them and retain a simpler baseline for comparison. If the same outcome can be achieved with fewer handoffs and clearer ownership, that option belongs in the release decision.

What to read next

References

  1. Galileo, Evaluate documentation
  2. LangSmith, Evaluation types
  3. OpenAI, AgentKit and expanded Evals