The Defender’s Guide to AI Agents·Chapter 4

The AI Agent Attack Chain

The attack chain

From here the comparison is retired. You have the shape; now you need the words people actually use in meetings, so the rest of this guide says agent rather than assistant, and names things as your vendors and your auditors will name them.

You already think in chains. Something gets in, it establishes itself, it moves, it takes what it came for. The instinct is right and the stages are different, so it is worth walking through them once, slowly, with the honest answer at each step about whether you could break it there.

Eight stages. The eighth - suppressing the evidence - is not in the AI threat taxonomies at all, and it is in the general adversary framework, where both major SIEM vendors already ship detections for it against AI platforms. It belongs in any chain that ends in an investigation.

Eight-stage AI agent attack chainThe attack progresses from placement, retrieval, interpretation, and expansion to action, egress, persistence, and evidence suppression. Defenders have the strongest control over the final four stages. STAGES 1-4 · IN CONTENT AND INSIDE THE MODEL STAGES 5-8 · IN YOUR SYSTEMS - WHERE THE DEFENSE LIVES 1 Placement 2 Retrieval PARTLY 3 Interpretation 4 Expansion 5 Action BREAKABLE 6 Egress BREAKABLE 7 Persistence BREAKABLE 8 Suppression BREAKABLE Stage 8 is absent from the AI threat taxonomies. It belongs in any chain that ends in an investigation.
Most security spending on AI today targets stage 3 - trying to detect the malicious instruction. That is the hardest stage to win and the least rewarding. The last four stages are ordinary engineering, and they are where the outcome is actually decided.
StageWhat happensWhat you would seeCan you break it here?
1 · Placement The attacker gets their text somewhere the agent will eventually read it: a web page, an email, a support ticket, a shared document, a product review, a code comment. Nothing. This happens outside your organization, and often outside anyone's. No. You do not control the internet, and much of your own inbound content is written by strangers by design.
2 · Retrieval The agent pulls that content in - because a user asked a question it seems to answer, or because it was in the mailbox, or because it was in the search index. A retrieval or search event, if you collect one. Usually you do not. Partly. You choose which collections of documents an agent is allowed to search. Narrowing that is one of the few genuinely effective moves available.
3 · Interpretation The model reads the planted text as an instruction. There is no vulnerability here and nothing is exploited; the system is working normally. Nothing. There is no event for "the model changed its mind." Not reliably. This is where most products are sold and where least is won. See Chapter 6, detection six.
4 · Expansion The agent gathers more than the task required, or asks a second agent - one with more access - to do something on its behalf. Search and tool events that are individually reasonable. The pattern is wrong; no single line is. Partly. Only if you recorded which agent called which, and where the original instruction came from.
5 · Action A tool is called, a record is changed, a message is drafted - or an answer is simply rendered to a screen, which as Chapter 1 showed is enough. The tool call. This is your highest-value event and the one most often not collected. Yes. What the agent is permitted to do, and under whose authority, is entirely yours to decide.
6 · Egress Something leaves: data in a request, an email, a file, a web fetch triggered by rendering. An outbound connection - from your network, or from the vendor's, which is the hard part. Yes. Where an agent may connect is a list you can write. Most organizations have never written it.
7 · Persistence Something is written to memory, a knowledge base, or a configuration, so the effect repeats without the attacker returning. A memory write. Almost nobody logs these. Yes. Rules about what may be written, by whom, and for how long are straightforward to impose - once you decide to.
8 · Suppression The evidence that would reveal any of the above is degraded or removed - tracing disabled, sampling reduced, retention shortened, logs deleted. A gap. Which is only visible if you were watching the telemetry itself. Yes, and it is the cheapest of all: treat loss of observability on a high-risk agent as an automatic reduction in its autonomy. See 5.11.

How this differs from the kill chain you already know

Close enough to be useful, different enough to mislead you if you do not notice where it diverges.

  • There is no exploit. Nothing is overflowed, escalated or bypassed. The step that corresponds to "exploitation" is the model reading a sentence, which is the thing it is for.
  • There is no malware, so there is nothing to find afterwards. No file on disk, no persistence in the registry, no sample to send to a lab. Your entire post-incident toolkit assumes an artifact exists.
  • There is usually no command and control. The attacker gives one instruction and leaves. They are not steering anything, so there is no beaconing to spot.
  • The chain can pause for weeks. Content placed on a Tuesday may not be retrieved until the following month, and memory written today may fire in April. Your investigation window is the wrong shape.
  • Delivery is content, not code. Everything you own for inspecting deliveries is looking for executable things.

Say this on Monday

"We can't stop the agent being told the wrong thing. We can decide what it's allowed to do afterwards, where it's allowed to connect, and what it's allowed to remember. Let's spend there."

About the author

Jessen Kurien is a cybersecurity leader and the author of The Defender’s Guide to AI Agents. His 18+ years in cybersecurity include nearly 15 years at Microsoft, work as part of the founding team of the Microsoft Threat Intelligence Center, and detection engineering leadership in Microsoft Defender XDR. His work connects investigations, detection engineering and security operations with the evidence and accountability needed for AI security and governance.

Meet Jessen

Connect with Jessen Speaking, workshops and training

This guide will go out of date.

Providers change how their logs work, models get retired, and new cases get disclosed. Ask to be told when this changes — no newsletter, just the updates.

Get told when it changes

Download the complete guide (PDF)

The telemetry contract, detection specifications, framework mappings and checklists are also published as files — the defender pack, CC BY 4.0, free to reuse.