The Defender’s Guide to AI Agents·Chapter 3

The Five Places an AI Agent Can Be Attacked

The five places an agent can be attacked

An agent is not one thing, and "we secured our AI" usually means one of these five was addressed and the other four were not. Each has its own way of being attacked, its own evidence, and its own fix.

Each one below ends with numbered steps you can actually follow this week. Where a step names a console or an export, Chapter 5 has the full platform-by-platform detail - what to switch on, what it gives you, and what it withholds.

Five AI agent attack surfacesAn AI assistant is surrounded by five attack surfaces: context, tools, memory, identity, and orchestration. THE ASSISTANT reads · decides · acts 1 · CONTEXT What goes in email, documents, pages 2 · TOOLS What it can do the buttons it may press 3 · MEMORY What it remembers believed again tomorrow 4 · IDENTITY Who it is its badge and its keys 5 · ORCHESTRATION Who it talks to other assistants
Five surfaces, five separate sets of controls. In most organizations today, only surfaces 2 and 4 are covered at all - and surfaces 1, 3 and 5 are where the documented attacks are landing.

Surface 1

Context - what goes in

What it is. Everything the assistant reads in order to do its job. Your instruction, plus every email, ticket, document, web page, calendar invite and search result it pulls in along the way.

How it is attacked. This is the one from Chapter 1. An attacker puts an instruction inside content they know the assistant will read. It can be invisible to a person - white text on white, a comment in a file's hidden properties, text sized at zero - and perfectly readable to the assistant, because the assistant is reading the file, not looking at the page.

In the wild. This is now the most widely documented attack class against agents, and the incident in Chapter 1 is one disclosed example of it against a shipping product used by millions. It needs no malware and no access - only a document, a page, or an email that your assistant will eventually read. Note that in that case the user never even opened the email; the assistant found it on their behalf.

What it looks like in a log

Almost nothing, by default. Reading a document is not an event most companies record. The visible part is what happens next: a tool call that does not fit the task, a sudden jump in the amount of text processed, or an outbound request to somewhere the assistant has never been.

One thing to do this week - build the untrusted-input list

  1. For each assistant, write down every place it reads from: mailboxes, document libraries, ticket queues, chat channels, the open web.
  2. Against each one, answer a single question: can someone outside this company put content there? Email is yes. A customer-facing ticket queue is yes. Web browsing is yes. That subset is your real attack surface.
  3. Do not trust the list. Go and read what it actually opened: in a Microsoft estate, Purview Audit records of type CopilotInteraction carry a field called AccessedResources listing every file and message it touched to answer. Compare that to your list.
  4. Turn off web access for any assistant whose job does not need it. In the same audit records, an AISystemPlugin entry naming the web search tool marks the answers that reached the internet.
  5. Hand the gap between your list and the log to whoever owns the assistant. That gap, not the list, is the finding.

Surface 2

Tools - what it can do

What it is. The buttons. Each tool is described to the assistant in plain language - "use this to send an email" - and the assistant chooses between them based on those descriptions.

How it is attacked. Two ways, and they are quite different. First, the description itself can carry a hidden instruction, because the assistant reads the description as text like everything else. Second, and more ordinary: the plug-in is software from a public repository, and software from public repositories gets updated, taken over and backdoored.

In the wild. In September 2025 an email plug-in named postmark-mcp was benign for fifteen versions and then, at version 1.0.16, quietly began copying every email it handled to an outside address. The vendor confirmed the pattern; how many organizations were affected is not known, and the figure of a few hundred that circulated at the time was an estimate derived from download counts rather than a victim count. Nobody had to break in. They had already been invited.

The wider picture is a question of exposure rather than proven holes. A 2026 review of 2,614 of these plug-ins found 82% using file operations that sit in the path-traversal risk class and 67% using APIs associated with code injection. Read that carefully: it means most of them are working in the neighborhoods where those bugs live, not that most of them are vulnerable. It is an argument for reviewing them, not for panic. Separately, one popular connector (mcp-remote, CVE-2025-6514) was rated 9.6 out of 10 by the researchers who found it, in a package with over 437,000 downloads - though exploiting it required the user to connect to a server the attacker controlled.

What it looks like in a log

The tool execution record: what the agent attempted, which system and target it acted on, the parameters used, and whether the action succeeded. This is the highest-value event you can collect and the one most often missing. A change in the description of a tool is a second event worth watching, and almost nobody watches it.

One thing to do this week - inventory the plug-ins

  1. Collect the configuration for every agent - the file or console screen listing its connected tools and MCP servers. Developers usually know where this lives; ask for it rather than looking for it.
  2. For each entry record five things: name, version, publisher, where it was downloaded from, and who approved it. Any blank is a finding on its own.
  3. Find out what is actually being called, not what is installed. On Claude Enterprise the Compliance API returns, for every tool call, the tool name, the integration name and the MCP server URL - which is the real inventory.
  4. Ask the one question that matters: what happens when one of these updates? If the answer is "it updates automatically," stop there and write it up.
  5. Pin every entry to a specific version, turn off auto-update, and make adding a new one a change with a person’s name against it.

Surface 3

Memory - what it remembers

What it is. Notes the assistant keeps so it does not start from zero every time. Preferences, facts about your systems, conclusions it reached last week.

How it is attacked. The attacker gets a false fact written into the notebook. Then they leave. Tomorrow, next week, or next month, the assistant reads its own notes and believes them - because they are its own notes, and nothing in the notebook records where each line came from.

This is the surface that breaks most people's intuition, for two reasons. The attack and the damage are separated in time, sometimes by weeks, so the investigation looks in the wrong fortnight. And the note can be read by a different assistant with more permissions than the one that wrote it, which means your careful access controls are walked around rather than broken.

In the wild. Demonstrated repeatedly by researchers. Detection is genuinely immature, and the clearest evidence of that comes from the strongest published attempt at it: a 2026 preprint reports an area-under-curve of 0.99 on its own test set, and then, in a follow-up on benign traffic, a false-positive rate between 24.7% and 52.6% depending on the protocol being watched. The author’s own conclusion is the useful part - the signal is a precondition for an attack, not evidence of one, and blocking on it is not viable. A preprint that argues against its own deployment is worth more than a product that does not.

What it looks like in a log

A memory write event: what was written, by which assistant, in which session, after reading what. Almost no organization collects this today. Without it, the delayed-detonation attack is not merely hard to find - there is literally no record that the write occurred.

One thing to do this week - find the shared memory

  1. Ask, per agent: does it keep anything between sessions? Saved preferences, project knowledge, a notes file, a vector store, a knowledge base. The answer is often yes and often surprises the person answering.
  2. For each store, list which agents can write to it and which can read it. Those two lists are rarely the same and almost never written down.
  3. Look for one specific pattern: a store written by an agent that reads outside content and read by an agent with higher permissions. That is a route around your access controls that leaves no trace.
  4. Check whether anything is recorded when a memory is written. On every major platform today the honest answer is no. Write that down as an accepted gap rather than discovering it during an investigation.
  5. As an interim control, agree that memories carry where they came from and that they expire. Both are cheap, and neither needs a product.

Surface 4

Identity - who it is

What it is. The account and the keys the assistant uses to reach your systems. This is the most familiar surface, and the one your existing skills transfer to almost completely.

How it is attacked. Conventionally, which is the good news. Keys get stolen, stored in the wrong place, shared between assistants, or given permissions nobody has reviewed. What is new is the scale and the speed: assistants get created quickly, often by developers, often without going near an identity process.

In the wild. A 2026 scan of MCP-related configuration files in public code repositories found 24,008 unique secrets sitting in plain text, of which 2,117 were still valid at the time of scanning. Those are files people published themselves, not servers anyone had to break into.

Stolen keys are attractive for three separate reasons, and it is worth separating them: the key is loot (it reaches data), it is compute (the attacker runs their own AI on your bill), and it is cover (their activity looks like your legitimate traffic).

What it looks like in a log

Your normal identity telemetry, which is the point - you already have this. What is usually missing is the link: knowing which account belongs to which assistant, and which human owns that assistant. Without that mapping, an identity alert tells you an account misbehaved and nothing about what to do next.

One thing to do this week - count the agent identities

  1. This is the one surface where the tooling is genuinely ahead of you, so start in the console. In a Microsoft estate, agent identities now appear in the directory as their own object type with their own listing screen, and through the directory API.
  2. Pull the sign-ins. Agent sign-ins carry their own event type and an agent-type field, so they can be filtered out of the noise rather than hunted for.
  3. Cross-check against a second source, because no single inventory is complete: the security portal keeps its own agent inventory, and in AWS the calling identity appears on every model invocation recorded in CloudTrail.
  4. Add two columns the tooling will not give you: which human owns this, and what does it actually reach - read from the live permissions, not from a design document.
  5. Anything with no owner gets an owner or gets disabled. Set an expiry date on every one, so the list has to be justified again rather than growing quietly.

Surface 5

Orchestration - who it talks to

What it is. One assistant asking another to do something. A research assistant hands findings to a writing assistant; a triage assistant asks a remediation assistant to act.

How it is attacked. By reaching the one you can get to, in order to reach the one you cannot. A low-privilege assistant that reads outside content becomes the way in to a high-privilege assistant that trusts it. The second assistant has no way of knowing that the first one is repeating something it was told by a stranger - because, again, it all arrives in the same voice.

In the wild. Researchers demonstrated one plug-in reaching across to another and pulling out message history through it, using a completely unrelated and apparently harmless connector as the entry point. The compromised component had no access to the data it ultimately exposed. The one it talked to did.

What it looks like in a log

Assistant-to-assistant calls, if you record them at all. The field that matters is the one nobody keeps: where did this instruction originally come from, three hops back. Without it you can see the conversation and not its origin.

One thing to do this week - draw the map

  1. One page. A box for every agent, an arrow for every agent that can call another. Whiteboard is fine; this does not need a tool.
  2. Mark every box that reads content from outside the organization. Mark every box that holds permissions you would care about losing.
  3. Find any arrow running from the first kind of box to the second. That arrow is the whole problem in one line, and it is where the main control in Chapter 7 goes.
  4. For each such arrow ask whether the receiving agent can tell where the instruction originally came from. The answer today is almost always no, because nothing carries the origin.
  5. Keep the page. It is the fastest briefing you will ever give an executive, and it will be out of date within a month - which is itself worth showing them.

The pattern across all five: these attacks are not hard to detect because they are subtle. They are hard to detect because the events that would show them are not written to any log you keep.

That is genuinely different from most of security, where the attacker is hiding among your data. Here there is nothing to hide among. Chapter 5 is about fixing that, and it is the most useful thing most teams can do this quarter.

Say this on Monday

"There are five places an agent can be attacked. We have controls on two of them. Let's start by finding out what's true for the other three."

SourcesLast verified 17 September 2026

  1. The postmark-mcp backdoor: vendor advisory, 25 September 2025. postmarkapp.com · malicious package record MAL-2025-47604, osv.dev. The disclosing firm’s original writeup is no longer reachable following an acquisition.
  2. CVE-2025-6514 in mcp-remote, affecting versions 0.0.5 to 0.1.15, patched in 0.1.16. Rated 9.6 by JFrog, who found and assigned it; NIST has not enriched the record. Exploitation requires connecting to a server the attacker controls. nvd.nist.gov · research.jfrog.com
  3. The 2,614-implementation figures come from a commercial dependency report, not a peer-reviewed study, and measure use of APIs in a risk class rather than confirmed vulnerability. Endor Labs, January 2026. endorlabs.com
  4. 24,008 unique secrets and 2,117 valid credentials in MCP-related configuration files in public repositories. GitGuardian, State of Secrets Sprawl 2026, 17 March 2026. blog.gitguardian.com
  5. Memory-poisoning detection: Leong, “Forensic Trajectory Signatures for Agent Memory Poisoning Detection,” arXiv preprint, June-July 2026. Not peer reviewed. AUC 0.9904; benign false-positive rate 24.7-52.6%; the author concludes standalone blocking is not viable. arxiv.org/abs/2606.30566

About the author

Jessen Kurien is a cybersecurity leader and the author of The Defender’s Guide to AI Agents. His 18+ years in cybersecurity include nearly 15 years at Microsoft, work as part of the founding team of the Microsoft Threat Intelligence Center, and detection engineering leadership in Microsoft Defender XDR. His work connects investigations, detection engineering and security operations with the evidence and accountability needed for AI security and governance.

Meet Jessen

Connect with Jessen Speaking, workshops and training

This guide will go out of date.

Providers change how their logs work, models get retired, and new cases get disclosed. Ask to be told when this changes — no newsletter, just the updates.

Get told when it changes

Download the complete guide (PDF)

The telemetry contract, detection specifications, framework mappings and checklists are also published as files — the defender pack, CC BY 4.0, free to reuse.