The Defender’s Guide to AI Agents·Chapter 6

AI Agent Detections: What to Look For

What to look for

6.1 · Six detections to prioritize. Start with controls your telemetry can support and your team can validate. The sequence below is a proposed starting point, not a measured ranking of detection quality. Establish a baseline, test representative benign and malicious cases, and adjust priorities to the risk in your environment. 6.2 gives the full catalog.

For each: the idea, likely sources of false alarms, and coverage limits. These are design considerations, not measured performance claims. Ask vendors for results from a representative evaluation and for the cases their controls miss.

Detection one · highest value

The model that answered is not the model you asked for

The idea. Where instrumentation records both the requested and returned model identifiers, compare them against an approved mapping of deployment names, aliases and model versions. OpenTelemetry defines gen_ai.request.model and gen_ai.response.model, but field availability depends on the integration. Investigate an unexpected mapping; track missing identifiers as a telemetry gap rather than treating absence as a match.

False alarms. Approved aliases, deployment-to-model mappings, fallback routing and version changes can produce legitimate differences. Maintain the mapping and measure alert precision on your own traffic before setting severity or enabling automated action.

What it catches. Silent version substitution, a forgotten floating version name in production code, and a provider changing what a name points at. Chapter 8 explains why this matters more than it sounds.

What it misses. Attacks that use the expected model, and changes hidden by incomplete or inaccurate telemetry. A mismatch is a change signal that needs investigation; it does not establish malicious activity.

Detection two

An agent called a tool it has never called before

The idea. Agents are creatures of habit. Learn the set of tools each one actually uses over two weeks, then alert on anything outside it.

False alarms. New tasks, approved tools and seasonal workflows can produce legitimate deviations. Baseline by agent role, review change records, and measure the resulting analyst workload.

What it catches. The moment an agent is redirected. Stage 5, at the point of action, which is where you want to be standing.

What it misses. An attack that stays entirely within the tools the agent already uses every day - which, for a read-and-answer agent, is most of them.

Detection three

The agent connected somewhere it has no business connecting

The idea. Write down where each agent is allowed to reach. Alert - or better, block - on anything else.

False alarms. Legitimate destinations change as integrations and business workflows evolve. Assign an owner to the allowlist, review exceptions, and validate coverage before relying on it for containment.

What it catches. Stage 6, the point where data actually leaves.

What it misses. Two important things. If the agent runs in a vendor's cloud, the connection may never touch your network, and you will see nothing. And an attacker who routes the data through a service already on your list - a file-sharing site, a form, the AI vendor's own domain - passes straight through. Narrowing the list aggressively is the main lever you have here, and it is a partial one.

Detection four

This session read far more than it needed to

The idea. Track how much content each agent takes in per session. Alert on sessions well outside the normal range, especially where a large share came from outside the organization.

False alarms. Large legitimate tasks can resemble excessive collection. Set thresholds by workload and agent role, then measure the false-positive rate and review burden on representative traffic.

What it catches. Stage 4, expansion - the agent gathering far more than the job required.

What it misses. The precise, surgical version, which is the good version. An instruction asking for one specific fact produces a session that looks entirely normal.

Detection five

A memory was written right after untrusted content was read

The idea. An ordering rule rather than a content rule. If an agent reads external content and then writes something to long-term memory in the same session, that is worth a look - particularly if the memory is one a more privileged agent will later read.

False alarms. Frequent legitimate memory updates may generate noise. Baseline normal write behavior and correlate writes with provenance, permissions and the surrounding task.

What it catches. Stage 7, the delayed-detonation case, at the only moment it is visible - weeks before it does anything.

What it misses. Delayed writes and activity outside the correlation window. This also depends on the memory-write telemetry described in Chapter 5; verify that your platform exposes it and that collection is complete.

Detection six · handle with care

Something in that text looked like an injection

The idea. Inspect incoming content, or the agent's own behavior, and classify it as an attempted manipulation. This is what most products on the market are.

False alarms. Performance depends on the detector, threshold, dataset and workload. A strong area-under-curve result does not establish an acceptable production operating point. Measure both the false-positive rate (benign cases incorrectly flagged divided by all benign cases) and alert precision (true positives divided by all positive alerts). They use different denominators; neither can be inferred from the other without the case mix and detection results.

What it catches. The clumsy attempts. Which is most attempts, so this is genuinely worth having.

What it misses. Paraphrased, indirect or context-dependent attacks may evade the detector. Evaluate those cases explicitly; a keyword list alone cannot establish whether an instruction is authorized.

How to buy this without being sold to. Ask for the false-positive rate on your traffic, not theirs. Ask what happens when the underlying model changes next month. And place it as an enrichment on an alert you already trust rather than as an alert of its own - it is a good second opinion and a poor first one.

Notice the shape: the reliable detections all sit at the end of the chain, and the unreliable one sits at the start.

That is not a coincidence, and it is the single most useful thing in this part. Detecting the bad instruction is a judgment about meaning, and meaning is contested. Detecting an unexpected tool call or an unexpected destination is a comparison against a list. Spend where the comparison is possible.

Say this on Monday

"Three of our six detections need no AI at all - they're lists and comparisons. Let's build those before we evaluate a single product."


The full catalog

Six is where you start. Thirty is roughly what a mature agent estate needs, and the useful way to read the list is by signal quality rather than by threat.

Deterministic

A comparison against an approved list or a recorded fact. The decision can be reproducible and easy to inspect, but its security value still depends on telemetry quality, baseline accuracy and context. Deterministic logic does not guarantee low false positives.

Behavioral

Needs a baseline first. Reliable once the baseline settles, noisy for the first fortnight, and sensitive to legitimate change.

Probabilistic

A judgement about meaning. Useful as enrichment on an alert you already trust. Dangerous as the sole trigger for anything destructive.

Prioritize by threat coverage, evidence quality and operational cost. Deterministic checks and probabilistic models serve different purposes; validate both against the use cases they are intended to protect.

#DetectionQualityWhat it joins, and what it catches
Identity and authority
1Token audience or resource mismatchDeterministicIdentity events against tool calls. A token minted for one resource being presented to another - the clearest sign that a credential has escaped its intended path.
2Delegation expansionDeterministicDelegation chain against the original grant. The scope at step four exceeds the scope at step one. Should be impossible; frequently is not.
3Approval-to-action mismatchDeterministicThe two digests from 5.9. What ran is not what was approved.
4Approval replay, and approval fatigueDeterministic
+ behavioral
Approval records over time. The same approval used twice, an approval used after expiry - and separately, an approver whose median decision time has fallen to under two seconds, which is a control failing quietly.
5Action with no accountable authorityDeterministicConsequential actions against approval and delegation records. A change with no human anywhere behind it. This is the one an auditor will ask for.
Context and retrieval
6Untrusted context entering a privileged workflowDeterministicContext admission trust class against the acting identity’s privilege. The single most valuable detection in this table, because it is the violation the architecture in Chapter 7 exists to prevent.
7Retrieval crossing a permission or tenant boundaryDeterministicRetrieval authorization against the requester’s entitlements. Catches the case in Chapter 1 where the answer was accurate, permitted, and should never have been surfaced.
8Policy lost during compactionDeterministicCompaction events against what remained. The window filled and the safety instructions were what got dropped.
9Session read volume well outside normalBehavioralContext admission totals per session, weighted by how much came from outside. Catches expansion; misses the surgical version.
10Embedding or index driftDeterministicIndex and embedding-model change events. Re-embedding a corpus silently changes every future answer, and almost nobody treats it as a change at all.
Memory
11Memory written after untrusted content was readDeterministicMemory write joined to the preceding context admission. An ordering rule, not a content judgement - which is why it works.
12Memory privilege crossoverDeterministicWrite identity against read identity on the same store. Written by something low-privilege, read by something high-privilege. A route around your access controls that leaves no other trace.
13Memory contradicting an authoritative sourceBehavioralMemory contents against the system of record. Expensive, imperfect, and the only thing that finds a poisoning that already succeeded.
Tools and supply chain
14Tool catalog or description driftDeterministicHash the description a tool presents to the agent; alert when it changes. The description is an instruction, so changing it changes behavior without touching your code.
15Tool name or origin collisionDeterministicTwo tools claiming the same name, or a familiar name from an unfamiliar server. Cheap, and catches a whole class of substitution.
16Declared risk differing from enforced behaviorDeterministicThe tool’s declared risk tier against what it actually did. A tool registered as read-only that wrote.
17Protocol, header or body mismatchDeterministicWhat the client sent against what the server received. Where an interposed component is rewriting requests.
18Tool call with no valid model or plan parentDeterministicTrace parentage. An action with no reasoning above it did not come from the agent; something else called the tool.
19Provenance failureDeterministicRunning tool versions against the approved manifest. Unsigned, unpinned, or updated since approval.
20Tool called that this agent has never calledBehavioralThe original detection two. Still one of the best.
Execution and egress
21Unexpected code or process executionBehavioralRuntime telemetry against the agent’s normal shape. An agent that has never spawned a shell spawning one.
22Secret accessed outside its expected actionDeterministicSecret lease events against the action in flight. A credential retrieved for a task that does not use it.
23Novel egress, and remote renderingDeterministicOutbound destinations against the allowlist - including content the agent’s own answer causes to be fetched, which is the mechanism in Chapter 1.
24Browser, clipboard, upload and download anomaliesBehavioralBrowser automation telemetry. Redirects to unexpected origins, clipboard access, files leaving through a rendered page.
25Ephemeral sandbox lifecycle gapsDeterministicSandbox created, used, destroyed - and whether the evidence survived the destruction.
Outcome and integrity
26Claimed success with no authoritative confirmationDeterministicThe join from 5.10. The tool says it worked and the target system has no record of it.
27Partial success, or a rollback that did notDeterministicDownstream records against intended state. Particularly for permission changes, where half a change is worse than none.
28Telemetry suppression and sequence gapsDeterministicEverything in 5.11. Tracing disabled, sampling changed, hash chain broken, clock jumped.
Behavior and scale
29Runaway loops and resource harvestingBehavioralStep counts, spend, rate-limit events. Both a cost control and a compromise signal, and the finance team will help you build it.
30Shadow and orphan agentsDeterministicIdentities acting that are not in the registry, and registered agents whose owner has left. The first list is your real inventory problem.
31Model route and version driftDeterministicThe original detection one. Requested model against served model.
32Guardrail denial followed by a semantic retryBehavioralPolicy decisions in sequence. Denied, rephrased, allowed. One of the clearest attack signatures available, and it needs no content inspection at all.
33Cross-agent and cross-plane cascadeBehavioralHandoff events against the origin field. One instruction fanning out across several agents and several systems.
34Injection scoringProbabilisticEnrichment only. Attach it to an alert you already trust. Never let it be the sole trigger for a destructive automated response - see 6.1, detection six, for the false-positive arithmetic.

Count the qualities. Twenty-one of these are deterministic.

These checks compare recorded values and approved lists without asking a model to interpret intent. They can be practical starting points where the required telemetry exists. Review the analytic content already available in your platforms, identify coverage gaps, and test proposed additions before promoting them into production.


What you can actually buy or download today

Two honest structural facts before you go shopping, because both change what is possible.

There is no agent log-source taxonomy. The main community detection-rule format has no category for AI, LLM or agent log sources. Rules that claim to be agent detections use an invented source name, which means they do not travel through standard tooling the way an endpoint or cloud rule does. You will be normalizing into your own schema whatever you buy.

The common event schema has no agent class either. As of its August 2026 release, AI is represented as a profile and a set of objects layered onto existing classes rather than as classes of its own. An agent tool call is modeled as API activity. That is workable, and it is not the same as first-class support.

What existsMaturityHonest assessment
Cloud control-plane detections for AI services Real, vendor-maintained The major SIEM vendors ship analytic content for managed model services - disabled invocation logging, cross-region inference abuse, token abuse, oversized prompts. This is genuine, supported, and the fastest thing to turn on. It watches the platform, not the agent.
Provider-side threat protection Real, generally available Cloud providers now offer AI threat alerting with optional prompt evidence in the alert. Useful, and confined to what that provider can see.
Community agent rule packs Alpha Public rule collections for agent behavior exist. The largest is single-maintainer, pre-release, and asserts adoption that we could not independently confirm. Read them for ideas; do not build a program on them.
Detections for the twenty-one deterministic rows above Yours to write Nobody sells them, because they depend on joins that only exist once you have built the correlation spine in 5.5. This is the work.

What to ask a vendor. Not “do you detect prompt injection” - everyone says yes. Ask which of the sixteen event families they ingest, whether they carry your correlation identifiers through, what their false-positive rate is on your traffic, and what happens to their detections when the underlying model changes next month. The fourth question is the one that separates a product from a demo.

SourcesLast verified 17 September 2026

  1. The distinction between the model requested and the model that answered is carried in the OpenTelemetry generative-AI attributes as gen_ai.request.model and gen_ai.response.model. Note these conventions have moved to a separate repository and are marked accordingly in the main registry. opentelemetry.io
  2. False-positive figures for injection and memory-poisoning detection: Leong, arXiv preprint, 2026. arxiv.org/abs/2606.30566
  3. No AI, LLM or agent category exists in the community detection-rule log-source taxonomy. sigmahq.io/docs/basics/log-sources
  4. Vendor-maintained analytic content for managed model services, including detection of deleted model-invocation logging. research.splunk.com · docs.datadoghq.com
  5. The common event schema models AI through a profile and objects on existing classes rather than through dedicated classes. github.com/ocsf/ocsf-schema
↑ Top

About the author

Jessen Kurien is a cybersecurity leader and the author of The Defender’s Guide to AI Agents. His 18+ years in cybersecurity include nearly 15 years at Microsoft, work as part of the founding team of the Microsoft Threat Intelligence Center, and detection engineering leadership in Microsoft Defender XDR. His work connects investigations, detection engineering and security operations with the evidence and accountability needed for AI security and governance.

Meet Jessen

Connect with Jessen Speaking, workshops and training

This guide will go out of date.

Providers change how their logs work, models get retired, and new cases get disclosed. Ask to be told when this changes — no newsletter, just the updates.

Get told when it changes

Download the complete guide (PDF)

The telemetry contract, detection specifications, framework mappings and checklists are also published as files — the defender pack, CC BY 4.0, free to reuse.