The Defender’s Guide to AI Agents·Chapter 11

Three AI Agent Investigations, End to End

Three investigations, end to end

Everything so far has been architecture. This part is what it feels like to use it - three incidents worked through from first alert to evidence package, showing which sources participate, which fields do the joining, and where the investigation would stop dead if one of them were missing.

Each walkthrough names the event families from 5.6 and the detections from 6.2, so you can read it as a test of your own coverage. Where a step depends on something most organizations do not collect, it says so.

Investigation one

Injected content reaches a privileged tool

What fires. Detection 6 - untrusted context entering a privileged workflow. A support-triage agent, which reads customer tickets, made a call to an entitlements tool inside the same trace as a context admission classed as external.

StepSourceWhat it gives youJoin
1Context admission (family 2)A ticket body entered the working set. Trust class: external. Fingerprint recorded, content not.trace_id
2Retrieval authorization (3)Nothing unusual. The agent searched the knowledge base under its own entitlements and got what it should.trace_id
3Model invocation (4)Requested model matches served model. No drift. Token count 4× the session median.trace_id, span_id
4Planning (6)Plan revised mid-task. Reason code present. No chain-of-thought needed - the structured record is enough to see the shape change.parent_span_id
5Policy decisions (5)One denial, then a rephrased request, then an allow. Detection 32. This is the moment the investigation becomes an incident.trace_id, sequence_no
6Tool execution (7)Entitlements tool called. Target: a group the agent has never touched. Detection 20.tool_call_id
7Identity (9)Which identity presented which token, and its scope at that moment.delegation_id
8Authoritative outcome (12)The directory’s own audit record. The membership change did apply.downstream_audit_uid
9Egress (14)No outbound connection outside the allowlist. Nothing left.trace_id

Containment decision. Revoke the agent’s delegated token - not just stop the process, because the token outlives it. Drop the agent to propose-only. Reverse the membership change, then verify the reversal in the directory’s own log rather than trusting the tool’s success message. Quarantine the originating ticket so it is not re-read on the next run, which is the step most often forgotten.

What made it solvable: steps 1, 5 and 8. Without context admission you cannot show the instruction came from outside. Without the policy sequence you have a strange tool call and no story. Without the downstream record you do not know whether anything actually changed. Those three are the ones most estates lack.

Evidence package

Tier 0 for the whole trace. Tier 1 for the ticket body, the denied and allowed requests, and the tool arguments. Tier 2 opened on declaration, with the ticket preserved before the support platform’s own retention expires it - that clock is usually shorter than your investigation.

Investigation two

A memory written in March fires in May

What fires. Detection 12 - memory privilege crossover. A finance-reporting agent read a note written nine weeks earlier by a research agent that reads the open web. Different agents, different privilege, same store.

This is the investigation that breaks conventional tooling, because the cause and the effect are separated by two months and your retention window probably is not.

WhenSourceWhat it gives youJoin
MarchContext admission (2)The research agent read a vendor page. External trust class. Fingerprint recorded.trace_id (March)
MarchMemory lifecycle (8)The write, with the preceding context admission recorded on it. This single field is what connects the two halves. Without it there is no investigation, only a coincidence.memory_id, prior_context_id
MarchIdentity (9)Written by the research agent’s identity. Low privilege. Entirely permitted.agent_id
MayMemory lifecycle (8)The read. Different agent, higher privilege.memory_id
MayPlanning (6)The plan cites the memory as a source. The reason code names it.trace_id (May)
MayTool execution (7)A report generated on a false premise and distributed internally.tool_call_id
MayAuthoritative outcome (12)The document platform’s record: created, shared, opened eleven times.downstream_audit_uid

Containment. Expire the memory. Then find every decision downstream of it - which requires querying by memory_id across two months of traces, and is the moment you discover whether your retention was long enough. Notify the eleven readers. Assess whether anything was decided on the report.

The uncomfortable part. There was no attacker action in May. The agent behaved perfectly, using a source it was entitled to use, which happened to be false. If your model of an incident requires adversary activity at the time of impact, this one is invisible to you.

What most organizations would find instead

Nothing. Memory writes are the least-collected family in 5.6, and the field linking a write to what was read before it does not exist in any platform’s default output. You add it, or this investigation is not available to you at any price.

Investigation three

An approved action is not the action that ran

What fires. Detection 3 - approval-to-action mismatch. The digest recorded at approval does not match the digest computed at execution. The action was blocked before it ran, which is the entire point of 5.9.

StepSourceWhat it gives youJoin
1Supply chain (15)A tool server updated eleven days ago. Version changed; the approved manifest did not. Detection 19. Nobody looked.tool_id, tool_version
2Supply chain (15)The tool’s description also changed. Detection 14. The text a tool presents to an agent is an instruction, and this one now asks for a broader target.tool_id, description_hash
3Approval (10)What the approver was shown: one mailbox, read access, four-hour expiry. approved_action_digest recorded.approval_id
4Tool execution (7)The call the tool actually assembled: the whole mailbox database, read-write.tool_call_id
5Policy (5)execution_action_digest computed. Mismatch. Execution refused.approval_id
6Authoritative outcome (12)Queried and confirmed: no change at the target. The refusal held.downstream_audit_uid

Why this is the most valuable of the three. Nothing happened. There is no damage to assess, no notification to make, no rollback to verify. The investigation is about a near miss, and it is only visible because the digest comparison is a precondition rather than an alert.

Take the same sequence without 5.9 in place: the approver approves one mailbox, the tool reads the database, the tool reports success, and the trail shows an approved action that completed normally. The incident would be discovered, if ever, from the other end - by whoever eventually receives the data.

Containment

Pin the tool back to the approved version. Suspend the agents that use it. Check whether the eleven days between the update and the detection contain other calls to that tool, and whether any of them completed. That question is answerable in one query if you carry tool_version on the tool-call event, and is a week of work if you do not.

One field decides each of these three.

The context admission that shows the instruction came from outside. The memory write that records what was read immediately before it. The approval digest that can be compared with what executed. Each is a single field, none is expensive, and none is collected by default anywhere.

That is the honest summary of this guide’s technical half: the hard part is not analysis, it is that the evidence was never written down.

Say this on Monday

“Pick our most privileged agent and walk one of these three investigations against what we actually collect. Wherever we stop, that's the next thing we build.”

About the author

Jessen Kurien is a cybersecurity leader and the author of The Defender’s Guide to AI Agents. His 18+ years in cybersecurity include nearly 15 years at Microsoft, work as part of the founding team of the Microsoft Threat Intelligence Center, and detection engineering leadership in Microsoft Defender XDR. His work connects investigations, detection engineering and security operations with the evidence and accountability needed for AI security and governance.

Meet Jessen

Connect with Jessen Speaking, workshops and training

This guide will go out of date.

Providers change how their logs work, models get retired, and new cases get disclosed. Ask to be told when this changes — no newsletter, just the updates.

Get told when it changes

Download the complete guide (PDF)

The telemetry contract, detection specifications, framework mappings and checklists are also published as files — the defender pack, CC BY 4.0, free to reuse.