How to Red-Team Agentic Workflows, Authority and Security
Agentic red teaming challenges the workflow with relevant hostile and manipulative conditions, then verifies whether evidence, permissions and authority remain intact.
Define the threat and safe behaviour
State the attacker or failure capability, accessible systems, target, likely consequence and expected safe response. A general request to break the agent produces anecdotes. A defined threat model produces reproducible evidence.
Attack instructions and knowledge
Test prompt injection in user messages, documents, web pages and retrieved records. Include false policy, hidden instructions, source impersonation and attempts to make external content redefine the agent's role.
Attack permissions and authority
Try direct requests, social engineering, split transactions, role confusion and missing approvers. Test values immediately below, at and above authority limits. Verify that the surrounding system enforces the boundary rather than relying only on a prompt.
Attack tools, memory and handovers
Test malicious parameters, unavailable tools, poisoned tool output, duplicate retries, cross-user memory and incomplete handovers. Confirm that one compromised component cannot silently expand another agent's authority.
Test exfiltration and privacy
Attempt to reveal personal data, credentials, system instructions, confidential records and information from another case. Evaluate both final outputs and tool calls. Use authorised synthetic or protected test environments.
Preserve failures as regression tests
Record the exact system version, case, environment, trace, grader and consequence. After remediation, rerun the failure and previously passed cases. Passing the tested threat model does not establish resistance to every attack.
Keep independent challenge and operating ownership
Development teams need rapid tests, while consequential systems benefit from challenge by people who did not design the behaviour or rubric. The accountable business owner still decides whether the residual risk is acceptable.
