The Defender’s Guide to AI Agents·Chapter 11
Three AI Agent Investigations, End to End
By Jessen Kurien · CISM, CISA, ISO/IEC 42001 Lead Implementer & Lead Auditor ·
Three investigations, end to end
Everything so far has been architecture. This part is what it feels like to use it - three incidents worked through from first alert to evidence package, showing which sources participate, which fields do the joining, and where the investigation would stop dead if one of them were missing.
Each walkthrough names the event families from 5.6 and the detections from 6.2, so you can read it as a test of your own coverage. Where a step depends on something most organizations do not collect, it says so.
Investigation one
Injected content reaches a privileged tool
What fires. Detection 6 - untrusted context entering a privileged workflow. A support-triage agent, which reads customer tickets, made a call to an entitlements tool inside the same trace as a context admission classed as external.
| Step | Source | What it gives you | Join |
|---|---|---|---|
| 1 | Context admission (family 2) | A ticket body entered the working set. Trust class: external. Fingerprint recorded, content not. | trace_id |
| 2 | Retrieval authorization (3) | Nothing unusual. The agent searched the knowledge base under its own entitlements and got what it should. | trace_id |
| 3 | Model invocation (4) | Requested model matches served model. No drift. Token count 4× the session median. | trace_id, span_id |
| 4 | Planning (6) | Plan revised mid-task. Reason code present. No chain-of-thought needed - the structured record is enough to see the shape change. | parent_span_id |
| 5 | Policy decisions (5) | One denial, then a rephrased request, then an allow. Detection 32. This is the moment the investigation becomes an incident. | trace_id, sequence_no |
| 6 | Tool execution (7) | Entitlements tool called. Target: a group the agent has never touched. Detection 20. | tool_call_id |
| 7 | Identity (9) | Which identity presented which token, and its scope at that moment. | delegation_id |
| 8 | Authoritative outcome (12) | The directory’s own audit record. The membership change did apply. | downstream_audit_uid |
| 9 | Egress (14) | No outbound connection outside the allowlist. Nothing left. | trace_id |
Containment decision. Revoke the agent’s delegated token - not just stop the process, because the token outlives it. Drop the agent to propose-only. Reverse the membership change, then verify the reversal in the directory’s own log rather than trusting the tool’s success message. Quarantine the originating ticket so it is not re-read on the next run, which is the step most often forgotten.
What made it solvable: steps 1, 5 and 8. Without context admission you cannot show the instruction came from outside. Without the policy sequence you have a strange tool call and no story. Without the downstream record you do not know whether anything actually changed. Those three are the ones most estates lack.
Evidence package
Tier 0 for the whole trace. Tier 1 for the ticket body, the denied and allowed requests, and the tool arguments. Tier 2 opened on declaration, with the ticket preserved before the support platform’s own retention expires it - that clock is usually shorter than your investigation.
Investigation two
A memory written in March fires in May
What fires. Detection 12 - memory privilege crossover. A finance-reporting agent read a note written nine weeks earlier by a research agent that reads the open web. Different agents, different privilege, same store.
This is the investigation that breaks conventional tooling, because the cause and the effect are separated by two months and your retention window probably is not.
| When | Source | What it gives you | Join |
|---|---|---|---|
| March | Context admission (2) | The research agent read a vendor page. External trust class. Fingerprint recorded. | trace_id (March) |
| March | Memory lifecycle (8) | The write, with the preceding context admission recorded on it. This single field is what connects the two halves. Without it there is no investigation, only a coincidence. | memory_id, prior_context_id |
| March | Identity (9) | Written by the research agent’s identity. Low privilege. Entirely permitted. | agent_id |
| May | Memory lifecycle (8) | The read. Different agent, higher privilege. | memory_id |
| May | Planning (6) | The plan cites the memory as a source. The reason code names it. | trace_id (May) |
| May | Tool execution (7) | A report generated on a false premise and distributed internally. | tool_call_id |
| May | Authoritative outcome (12) | The document platform’s record: created, shared, opened eleven times. | downstream_audit_uid |
Containment. Expire the memory. Then find every decision downstream of it - which requires querying by memory_id across two months of traces, and is the moment you discover whether your retention was long enough. Notify the eleven readers. Assess whether anything was decided on the report.
The uncomfortable part. There was no attacker action in May. The agent behaved perfectly, using a source it was entitled to use, which happened to be false. If your model of an incident requires adversary activity at the time of impact, this one is invisible to you.
What most organizations would find instead
Nothing. Memory writes are the least-collected family in 5.6, and the field linking a write to what was read before it does not exist in any platform’s default output. You add it, or this investigation is not available to you at any price.
Investigation three
An approved action is not the action that ran
What fires. Detection 3 - approval-to-action mismatch. The digest recorded at approval does not match the digest computed at execution. The action was blocked before it ran, which is the entire point of 5.9.
| Step | Source | What it gives you | Join |
|---|---|---|---|
| 1 | Supply chain (15) | A tool server updated eleven days ago. Version changed; the approved manifest did not. Detection 19. Nobody looked. | tool_id, tool_version |
| 2 | Supply chain (15) | The tool’s description also changed. Detection 14. The text a tool presents to an agent is an instruction, and this one now asks for a broader target. | tool_id, description_hash |
| 3 | Approval (10) | What the approver was shown: one mailbox, read access, four-hour expiry. approved_action_digest recorded. | approval_id |
| 4 | Tool execution (7) | The call the tool actually assembled: the whole mailbox database, read-write. | tool_call_id |
| 5 | Policy (5) | execution_action_digest computed. Mismatch. Execution refused. | approval_id |
| 6 | Authoritative outcome (12) | Queried and confirmed: no change at the target. The refusal held. | downstream_audit_uid |
Why this is the most valuable of the three. Nothing happened. There is no damage to assess, no notification to make, no rollback to verify. The investigation is about a near miss, and it is only visible because the digest comparison is a precondition rather than an alert.
Take the same sequence without 5.9 in place: the approver approves one mailbox, the tool reads the database, the tool reports success, and the trail shows an approved action that completed normally. The incident would be discovered, if ever, from the other end - by whoever eventually receives the data.
Containment
Pin the tool back to the approved version. Suspend the agents that use it. Check whether the eleven days between the update and the detection contain other calls to that tool, and whether any of them completed. That question is answerable in one query if you carry tool_version on the tool-call event, and is a week of work if you do not.
One field decides each of these three.
The context admission that shows the instruction came from outside. The memory write that records what was read immediately before it. The approval digest that can be compared with what executed. Each is a single field, none is expensive, and none is collected by default anywhere.
That is the honest summary of this guide’s technical half: the hard part is not analysis, it is that the evidence was never written down.
Say this on Monday
“Pick our most privileged agent and walk one of these three investigations against what we actually collect. Wherever we stop, that's the next thing we build.”
About the author
Jessen Kurien is a cybersecurity leader and the author of The Defender’s Guide to AI Agents. His 18+ years in cybersecurity include nearly 15 years at Microsoft, work as part of the founding team of the Microsoft Threat Intelligence Center, and detection engineering leadership in Microsoft Defender XDR. His work connects investigations, detection engineering and security operations with the evidence and accountability needed for AI security and governance.
This guide will go out of date.
Providers change how their logs work, models get retired, and new cases get disclosed. Ask to be told when this changes — no newsletter, just the updates.
Download the complete guide (PDF)
The telemetry contract, detection specifications, framework mappings and checklists are also published as files — the defender pack, CC BY 4.0, free to reuse.