The Defender’s Guide to AI Agents·Chapter 12
Containing a Compromised AI Agent
By Jessen Kurien · CISM, CISA, ISO/IEC 42001 Lead Implementer & Lead Auditor ·
Response: containing authority, not just execution
Killing the process does not contain the agent.
Its identity still exists. Its delegated tokens are still valid until they expire. Its memory still holds whatever was written there. The tools it registered are still registered, the knowledge sources it poisoned are still poisoned, and any child agents it spawned are still running.
Everything an incident responder knows about containing a compromised host applies. It just is not sufficient, because the thing you are containing is an authority, and authority does not live in the process.
Nine playbooks. Each names the evidence that disappears first, what triggers a kill switch, what to verify downstream, which identities are implicated, what rollback can and cannot reach, what residual risk remains afterwards, and - the field most often blank - who owns the decision.
The containment ladder
Before the individual playbooks, the four levers. Most teams have the first and treat it as the whole set.
| Lever | What it stops | What it does not stop |
|---|---|---|
| Stop execution | The running task. Immediate, visible, reversible. | Anything already delegated. The agent restarts and resumes. |
| Reduce autonomy | Drops the agent to propose-only, or suspends its privileged tools. Keeps the business running. | Tokens already issued, and work already queued. |
| Contain the identity | Disable the agent identity so it can no longer obtain tokens; revoke what is outstanding. This is the one that actually contains. | Sessions already established, until they expire. Ask your platform how long that is, before the incident. |
| Sever the surface | Deregister the tool, expire the memory, quarantine the source, revoke the secret lease. | Effects already propagated into systems of record. |
Identity containment as a distinct capability from process termination is now first-party documented on at least one major identity platform, which lets you disable an agent identity while keeping it and its metadata for the investigation. Two things worth confirming in your own environment before you need them: whether disabling revokes outstanding refresh tokens, and how long revocation takes to propagate. Neither is always documented.
The nine playbooks
| Incident | Volatile evidence - take first | Kill switch when | Verify downstream | Residual risk after containment |
|---|---|---|---|---|
| Context or prompt injection | The admitted content before its source system ages it out; the in-flight context window; the policy decision sequence | A privileged tool was reached in the same trace as external content | Every tool call in the trace, against each target’s own audit record | The source is still readable and will be read again unless quarantined. Other agents may read the same source. |
| Delegation or token compromise | Token issuance and exchange records; the delegation chain; every resource that accepted the token | Token presented to a resource outside its audience, or scope expanded along the chain | Each resource server’s own log - the agent’s logs will not show what a stolen token did elsewhere | Sessions established before revocation. Anything the token was exchanged for. |
| Memory poisoning | The memory record with its write provenance; the reads since; the plans that cited it | A privileged agent read a store written by an agent that reads untrusted content | Every decision downstream of the read, by memory_id across your full retention |
Copies. Summaries derived from it. Human decisions already taken on it. |
| Tool or supply-chain compromise | The tool version and description hash as they were; the server identity; recent calls | Version or description changed outside a change record | Every call to that tool since the change - not just the one that alerted | Other agents using the same tool. Data already returned by it. |
| Unexpected code execution | Sandbox contents before destruction - usually seconds; process tree; files written; network from the sandbox | An agent that has never executed code executed code | Anything the code touched, in that system’s own log | Artifacts persisted outside the sandbox. Credentials the sandbox could reach. |
| Data exfiltration | Egress records; rendered remote content; upload and download events; what was in context | Any egress outside the allowlist, including fetches triggered by rendering the agent’s own answer | Volume against the destination; what the target system says was read | It has already left. Containment limits the next one, not this one. |
| Runaway or cascading agents | Handoff records with the origin field; step counts; spend | Step or spend thresholds, or a fan-out across agents from one origin | Every action each participating agent completed | Queued and durable tasks resume after restart unless explicitly drained. |
| Model or configuration drift | The configuration fingerprint in force; requested versus served model; eval results | Eval below threshold, or served model differs from requested | Consequential actions taken under the changed configuration | Everything decided during the drift window, which is usually longer than you think. |
| Telemetry failure | Collector state; configuration change records; the last verifiable hash-chain link | Observability lost on a high-risk agent - automatic, per 5.11 | Whether actions continued during the blind window, from the target systems | The blind window is unreconstructable. Say so in the report rather than implying coverage. |
Three fields every playbook needs and most omit
- What rollback cannot reach. An email sent, a message posted, a file a person already downloaded, a decision someone already acted on. Write the boundary down before the incident, because it is the thing the business will ask about first and it is not a technical answer.
- Who owns the decision. Not the security team - the business owner of the process the agent performs. If an agent has no named owner it should not have been running; if it is running anyway, the incident is the moment that becomes everyone’s problem at once.
- What the approver actually saw. Preserved as evidence, separately from what the agent did. In any incident involving a human approval, the difference between those two is the whole question.
Say this on Monday
“If we had to contain our most privileged agent this afternoon - who disables the identity, how long until its tokens stop working, and who signs off on stopping the business process it runs?”
For the supporting investigation records, see what to log for AI agents. For the permissions and ownership behind those actions, see AI agent identity security.
SourcesLast verified 17 September 2026
- Disabling an agent identity so it can no longer obtain tokens, while retaining the identity and its metadata - identity containment as distinct from stopping a process. learn.microsoft.com - disable agent identities
- Token revocation and cross-party session revocation primitives. RFC 7009 · OpenID CAEP 1.0
- Joint international guidance on adopting agentic AI services, including quarantining agent requests to delete logs and restricting agent permissions automatically on unexpected behavior. cisa.gov - Careful Adoption of Agentic AI Services
- Order of volatility and evidence handling - the generic principles underneath every playbook above. RFC 3227 · NIST SP 800-61r3
- Agent identity lifecycle, just-in-time scoped credentials and revocation propagation across protocol adapters. Cloud Security Alliance - Agentic AI IAM
About the author
Jessen Kurien is a cybersecurity leader and the author of The Defender’s Guide to AI Agents. His 18+ years in cybersecurity include nearly 15 years at Microsoft, work as part of the founding team of the Microsoft Threat Intelligence Center, and detection engineering leadership in Microsoft Defender XDR. His work connects investigations, detection engineering and security operations with the evidence and accountability needed for AI security and governance.
This guide will go out of date.
Providers change how their logs work, models get retired, and new cases get disclosed. Ask to be told when this changes — no newsletter, just the updates.
Download the complete guide (PDF)
The telemetry contract, detection specifications, framework mappings and checklists are also published as files — the defender pack, CC BY 4.0, free to reuse.