The Defender’s Guide to AI Agents·Chapter 6
AI Agent Detections: What to Look For
By Jessen Kurien · CISM, CISA, ISO/IEC 42001 Lead Implementer & Lead Auditor ·
What to look for
6.1 · Six detections to prioritize. Start with controls your telemetry can support and your team can validate. The sequence below is a proposed starting point, not a measured ranking of detection quality. Establish a baseline, test representative benign and malicious cases, and adjust priorities to the risk in your environment. 6.2 gives the full catalog.
For each: the idea, likely sources of false alarms, and coverage limits. These are design considerations, not measured performance claims. Ask vendors for results from a representative evaluation and for the cases their controls miss.
Detection one · highest value
The model that answered is not the model you asked for
The idea. Where instrumentation records both the requested and returned model identifiers, compare them against an approved mapping of deployment names, aliases and model versions. OpenTelemetry defines gen_ai.request.model and gen_ai.response.model, but field availability depends on the integration. Investigate an unexpected mapping; track missing identifiers as a telemetry gap rather than treating absence as a match.
False alarms. Approved aliases, deployment-to-model mappings, fallback routing and version changes can produce legitimate differences. Maintain the mapping and measure alert precision on your own traffic before setting severity or enabling automated action.
What it catches. Silent version substitution, a forgotten floating version name in production code, and a provider changing what a name points at. Chapter 8 explains why this matters more than it sounds.
What it misses. Attacks that use the expected model, and changes hidden by incomplete or inaccurate telemetry. A mismatch is a change signal that needs investigation; it does not establish malicious activity.
Detection two
An agent called a tool it has never called before
The idea. Agents are creatures of habit. Learn the set of tools each one actually uses over two weeks, then alert on anything outside it.
False alarms. New tasks, approved tools and seasonal workflows can produce legitimate deviations. Baseline by agent role, review change records, and measure the resulting analyst workload.
What it catches. The moment an agent is redirected. Stage 5, at the point of action, which is where you want to be standing.
What it misses. An attack that stays entirely within the tools the agent already uses every day - which, for a read-and-answer agent, is most of them.
Detection three
The agent connected somewhere it has no business connecting
The idea. Write down where each agent is allowed to reach. Alert - or better, block - on anything else.
False alarms. Legitimate destinations change as integrations and business workflows evolve. Assign an owner to the allowlist, review exceptions, and validate coverage before relying on it for containment.
What it catches. Stage 6, the point where data actually leaves.
What it misses. Two important things. If the agent runs in a vendor's cloud, the connection may never touch your network, and you will see nothing. And an attacker who routes the data through a service already on your list - a file-sharing site, a form, the AI vendor's own domain - passes straight through. Narrowing the list aggressively is the main lever you have here, and it is a partial one.
Detection four
This session read far more than it needed to
The idea. Track how much content each agent takes in per session. Alert on sessions well outside the normal range, especially where a large share came from outside the organization.
False alarms. Large legitimate tasks can resemble excessive collection. Set thresholds by workload and agent role, then measure the false-positive rate and review burden on representative traffic.
What it catches. Stage 4, expansion - the agent gathering far more than the job required.
What it misses. The precise, surgical version, which is the good version. An instruction asking for one specific fact produces a session that looks entirely normal.
Detection five
A memory was written right after untrusted content was read
The idea. An ordering rule rather than a content rule. If an agent reads external content and then writes something to long-term memory in the same session, that is worth a look - particularly if the memory is one a more privileged agent will later read.
False alarms. Frequent legitimate memory updates may generate noise. Baseline normal write behavior and correlate writes with provenance, permissions and the surrounding task.
What it catches. Stage 7, the delayed-detonation case, at the only moment it is visible - weeks before it does anything.
What it misses. Delayed writes and activity outside the correlation window. This also depends on the memory-write telemetry described in Chapter 5; verify that your platform exposes it and that collection is complete.
Detection six · handle with care
Something in that text looked like an injection
The idea. Inspect incoming content, or the agent's own behavior, and classify it as an attempted manipulation. This is what most products on the market are.
False alarms. Performance depends on the detector, threshold, dataset and workload. A strong area-under-curve result does not establish an acceptable production operating point. Measure both the false-positive rate (benign cases incorrectly flagged divided by all benign cases) and alert precision (true positives divided by all positive alerts). They use different denominators; neither can be inferred from the other without the case mix and detection results.
What it catches. The clumsy attempts. Which is most attempts, so this is genuinely worth having.
What it misses. Paraphrased, indirect or context-dependent attacks may evade the detector. Evaluate those cases explicitly; a keyword list alone cannot establish whether an instruction is authorized.
How to buy this without being sold to. Ask for the false-positive rate on your traffic, not theirs. Ask what happens when the underlying model changes next month. And place it as an enrichment on an alert you already trust rather than as an alert of its own - it is a good second opinion and a poor first one.
Notice the shape: the reliable detections all sit at the end of the chain, and the unreliable one sits at the start.
That is not a coincidence, and it is the single most useful thing in this part. Detecting the bad instruction is a judgment about meaning, and meaning is contested. Detecting an unexpected tool call or an unexpected destination is a comparison against a list. Spend where the comparison is possible.
Say this on Monday
"Three of our six detections need no AI at all - they're lists and comparisons. Let's build those before we evaluate a single product."
The full catalog
Six is where you start. Thirty is roughly what a mature agent estate needs, and the useful way to read the list is by signal quality rather than by threat.
Deterministic
A comparison against an approved list or a recorded fact. The decision can be reproducible and easy to inspect, but its security value still depends on telemetry quality, baseline accuracy and context. Deterministic logic does not guarantee low false positives.
Behavioral
Needs a baseline first. Reliable once the baseline settles, noisy for the first fortnight, and sensitive to legitimate change.
Probabilistic
A judgement about meaning. Useful as enrichment on an alert you already trust. Dangerous as the sole trigger for anything destructive.
Prioritize by threat coverage, evidence quality and operational cost. Deterministic checks and probabilistic models serve different purposes; validate both against the use cases they are intended to protect.
| # | Detection | Quality | What it joins, and what it catches |
|---|---|---|---|
| Identity and authority | |||
| 1 | Token audience or resource mismatch | Deterministic | Identity events against tool calls. A token minted for one resource being presented to another - the clearest sign that a credential has escaped its intended path. |
| 2 | Delegation expansion | Deterministic | Delegation chain against the original grant. The scope at step four exceeds the scope at step one. Should be impossible; frequently is not. |
| 3 | Approval-to-action mismatch | Deterministic | The two digests from 5.9. What ran is not what was approved. |
| 4 | Approval replay, and approval fatigue | Deterministic + behavioral | Approval records over time. The same approval used twice, an approval used after expiry - and separately, an approver whose median decision time has fallen to under two seconds, which is a control failing quietly. |
| 5 | Action with no accountable authority | Deterministic | Consequential actions against approval and delegation records. A change with no human anywhere behind it. This is the one an auditor will ask for. |
| Context and retrieval | |||
| 6 | Untrusted context entering a privileged workflow | Deterministic | Context admission trust class against the acting identity’s privilege. The single most valuable detection in this table, because it is the violation the architecture in Chapter 7 exists to prevent. |
| 7 | Retrieval crossing a permission or tenant boundary | Deterministic | Retrieval authorization against the requester’s entitlements. Catches the case in Chapter 1 where the answer was accurate, permitted, and should never have been surfaced. |
| 8 | Policy lost during compaction | Deterministic | Compaction events against what remained. The window filled and the safety instructions were what got dropped. |
| 9 | Session read volume well outside normal | Behavioral | Context admission totals per session, weighted by how much came from outside. Catches expansion; misses the surgical version. |
| 10 | Embedding or index drift | Deterministic | Index and embedding-model change events. Re-embedding a corpus silently changes every future answer, and almost nobody treats it as a change at all. |
| Memory | |||
| 11 | Memory written after untrusted content was read | Deterministic | Memory write joined to the preceding context admission. An ordering rule, not a content judgement - which is why it works. |
| 12 | Memory privilege crossover | Deterministic | Write identity against read identity on the same store. Written by something low-privilege, read by something high-privilege. A route around your access controls that leaves no other trace. |
| 13 | Memory contradicting an authoritative source | Behavioral | Memory contents against the system of record. Expensive, imperfect, and the only thing that finds a poisoning that already succeeded. |
| Tools and supply chain | |||
| 14 | Tool catalog or description drift | Deterministic | Hash the description a tool presents to the agent; alert when it changes. The description is an instruction, so changing it changes behavior without touching your code. |
| 15 | Tool name or origin collision | Deterministic | Two tools claiming the same name, or a familiar name from an unfamiliar server. Cheap, and catches a whole class of substitution. |
| 16 | Declared risk differing from enforced behavior | Deterministic | The tool’s declared risk tier against what it actually did. A tool registered as read-only that wrote. |
| 17 | Protocol, header or body mismatch | Deterministic | What the client sent against what the server received. Where an interposed component is rewriting requests. |
| 18 | Tool call with no valid model or plan parent | Deterministic | Trace parentage. An action with no reasoning above it did not come from the agent; something else called the tool. |
| 19 | Provenance failure | Deterministic | Running tool versions against the approved manifest. Unsigned, unpinned, or updated since approval. |
| 20 | Tool called that this agent has never called | Behavioral | The original detection two. Still one of the best. |
| Execution and egress | |||
| 21 | Unexpected code or process execution | Behavioral | Runtime telemetry against the agent’s normal shape. An agent that has never spawned a shell spawning one. |
| 22 | Secret accessed outside its expected action | Deterministic | Secret lease events against the action in flight. A credential retrieved for a task that does not use it. |
| 23 | Novel egress, and remote rendering | Deterministic | Outbound destinations against the allowlist - including content the agent’s own answer causes to be fetched, which is the mechanism in Chapter 1. |
| 24 | Browser, clipboard, upload and download anomalies | Behavioral | Browser automation telemetry. Redirects to unexpected origins, clipboard access, files leaving through a rendered page. |
| 25 | Ephemeral sandbox lifecycle gaps | Deterministic | Sandbox created, used, destroyed - and whether the evidence survived the destruction. |
| Outcome and integrity | |||
| 26 | Claimed success with no authoritative confirmation | Deterministic | The join from 5.10. The tool says it worked and the target system has no record of it. |
| 27 | Partial success, or a rollback that did not | Deterministic | Downstream records against intended state. Particularly for permission changes, where half a change is worse than none. |
| 28 | Telemetry suppression and sequence gaps | Deterministic | Everything in 5.11. Tracing disabled, sampling changed, hash chain broken, clock jumped. |
| Behavior and scale | |||
| 29 | Runaway loops and resource harvesting | Behavioral | Step counts, spend, rate-limit events. Both a cost control and a compromise signal, and the finance team will help you build it. |
| 30 | Shadow and orphan agents | Deterministic | Identities acting that are not in the registry, and registered agents whose owner has left. The first list is your real inventory problem. |
| 31 | Model route and version drift | Deterministic | The original detection one. Requested model against served model. |
| 32 | Guardrail denial followed by a semantic retry | Behavioral | Policy decisions in sequence. Denied, rephrased, allowed. One of the clearest attack signatures available, and it needs no content inspection at all. |
| 33 | Cross-agent and cross-plane cascade | Behavioral | Handoff events against the origin field. One instruction fanning out across several agents and several systems. |
| 34 | Injection scoring | Probabilistic | Enrichment only. Attach it to an alert you already trust. Never let it be the sole trigger for a destructive automated response - see 6.1, detection six, for the false-positive arithmetic. |
Count the qualities. Twenty-one of these are deterministic.
These checks compare recorded values and approved lists without asking a model to interpret intent. They can be practical starting points where the required telemetry exists. Review the analytic content already available in your platforms, identify coverage gaps, and test proposed additions before promoting them into production.
What you can actually buy or download today
Two honest structural facts before you go shopping, because both change what is possible.
There is no agent log-source taxonomy. The main community detection-rule format has no category for AI, LLM or agent log sources. Rules that claim to be agent detections use an invented source name, which means they do not travel through standard tooling the way an endpoint or cloud rule does. You will be normalizing into your own schema whatever you buy.
The common event schema has no agent class either. As of its August 2026 release, AI is represented as a profile and a set of objects layered onto existing classes rather than as classes of its own. An agent tool call is modeled as API activity. That is workable, and it is not the same as first-class support.
| What exists | Maturity | Honest assessment |
|---|---|---|
| Cloud control-plane detections for AI services | Real, vendor-maintained | The major SIEM vendors ship analytic content for managed model services - disabled invocation logging, cross-region inference abuse, token abuse, oversized prompts. This is genuine, supported, and the fastest thing to turn on. It watches the platform, not the agent. |
| Provider-side threat protection | Real, generally available | Cloud providers now offer AI threat alerting with optional prompt evidence in the alert. Useful, and confined to what that provider can see. |
| Community agent rule packs | Alpha | Public rule collections for agent behavior exist. The largest is single-maintainer, pre-release, and asserts adoption that we could not independently confirm. Read them for ideas; do not build a program on them. |
| Detections for the twenty-one deterministic rows above | Yours to write | Nobody sells them, because they depend on joins that only exist once you have built the correlation spine in 5.5. This is the work. |
What to ask a vendor. Not “do you detect prompt injection” - everyone says yes. Ask which of the sixteen event families they ingest, whether they carry your correlation identifiers through, what their false-positive rate is on your traffic, and what happens to their detections when the underlying model changes next month. The fourth question is the one that separates a product from a demo.
SourcesLast verified 17 September 2026
- The distinction between the model requested and the model that answered is carried in the OpenTelemetry generative-AI attributes as
gen_ai.request.modelandgen_ai.response.model. Note these conventions have moved to a separate repository and are marked accordingly in the main registry. opentelemetry.io - False-positive figures for injection and memory-poisoning detection: Leong, arXiv preprint, 2026. arxiv.org/abs/2606.30566
- No AI, LLM or agent category exists in the community detection-rule log-source taxonomy. sigmahq.io/docs/basics/log-sources
- Vendor-maintained analytic content for managed model services, including detection of deleted model-invocation logging. research.splunk.com · docs.datadoghq.com
- The common event schema models AI through a profile and objects on existing classes rather than through dedicated classes. github.com/ocsf/ocsf-schema
About the author
Jessen Kurien is a cybersecurity leader and the author of The Defender’s Guide to AI Agents. His 18+ years in cybersecurity include nearly 15 years at Microsoft, work as part of the founding team of the Microsoft Threat Intelligence Center, and detection engineering leadership in Microsoft Defender XDR. His work connects investigations, detection engineering and security operations with the evidence and accountability needed for AI security and governance.
This guide will go out of date.
Providers change how their logs work, models get retired, and new cases get disclosed. Ask to be told when this changes — no newsletter, just the updates.
Download the complete guide (PDF)
The telemetry contract, detection specifications, framework mappings and checklists are also published as files — the defender pack, CC BY 4.0, free to reuse.