The Defender’s Guide to AI Agents·Chapter 9
ISO 42001 for Organizations That Use AI but Did Not Build It
By Jessen Kurien · CISM, CISA, ISO/IEC 42001 Lead Implementer & Lead Auditor ·
Where governance fits
There is an international standard for managing AI - ISO/IEC 42001. It is worth understanding even if you never go near certification, because it gives you language your board already accepts and a structure for the arguments in this guide.
It is also worth understanding its limits, which are real, and are the reason this part sits at the end rather than the beginning.
You are an "AI user," and that is a real category
The most useful thing the standard does for most organizations is in its opening clauses, not its controls. It recognizes different roles in the AI world - the people who build models, the people who supply them, and the people who merely use them - and it applies to all three.
If your company buys AI rather than builds it, you are an AI user, and declaring that is the move that makes everything else tractable. You are not being asked to certify a model you have never seen. You are being asked to show that you know where your AI is, what it can reach, who owns it, and what you would do when it goes wrong - which is precisely this guide.
The controls that carry the weight
The standard has 38 controls. Roughly a dozen matter for an AI user, and the mapping below is deliberately short.
| Reference | What it is called | What it actually asks of you |
|---|---|---|
| A.9.2 / A.9.3 / A.9.4 | Use of AI systems | The cluster written for people in your position. A.9.4, intended use, is where you write down that this agent does X and not Y - which doubles as the boundary an injection is trying to cross. |
| A.6.2.8 | AI system event logs | Keep logs. The standard does not say which events - so Chapter 5 of this guide is your answer to it, and a good one. |
| A.6.2.6 | Operation and monitoring | Where your eval set from Chapter 8 lives. Monitoring means noticing it got worse, not only noticing it stopped. |
| A.6.2.5 | Deployment | Your reversibility phasing from Chapter 10 answers this directly. |
| A.10.2 / A.10.3 | Responsibilities and suppliers | Where the model-churn problem becomes contractual: notice periods, change notification, the right to pin a version. |
| A.4.4 | Tooling resources | Your inventory of plug-ins and tools has a home here, which surprises people. |
| A.3.2 | Roles and responsibilities | Every agent has a named human owner, or it should not exist. This is control two from Chapter 7 in the standard's own words. |
| A.2.2 | AI policy | The document everyone starts with and the one that matters least on its own. |
| Clause 6.1.4 | AI system impact assessment | A clause rather than a control. Think through what happens to real people when this is wrong, before it is. |
| Clause 6.3 | Planning of changes | Where the three change classes from Chapter 8 belong. |
One thing people look for and do not find: there is no single control called "human oversight." It is spread across A.9.2, A.6.2.6 and Clause 6.1.4. Knowing that saves an afternoon and a certain amount of embarrassment.
What the standard gives you, and what it does not
What it gives you. A structure, a vocabulary your executives and your customers already recognize, a reason to fund the inventory, and a named owner for every agent. That is genuinely valuable and it is most of the value.
What it does not give you. Any of the technical content of this guide. A.6.2.8 says keep event logs; it does not say which six. Nothing in the standard tells you to split the reading identity from the acting identity, to pin a model version, or to watch for a memory write after untrusted content. The standard is the management system. The controls are still yours to design.
A certificate says an organization has a system for managing this. It does not say the system is any good.
That is not a criticism of the standard - no management system standard claims otherwise. It matters because the people selling certification will not always draw the line clearly, and because a defender who assumes the certificate covers the technical ground will not build any of it.
What an auditor will ask, and what usually happens
Useful whichever side of the table you are on. These are the questions where the honest answer is currently uncomfortable in most organizations.
| The question | What tends to be produced, and why it is not the answer |
|---|---|
| "Show me your AI system event logs." | A chat transcript. But a transcript shows what was typed, not what the agent retrieved, which tools it called, or which model version answered. It is the wrong artifact, and an auditor without a technical background will usually accept it. |
| "Why did the agent do this, in this case, last quarter?" | The output and the approval. Not the reasoning, not the content that prompted it, and often not the model version - which by then no longer exists to be asked. |
| "What is this agent permitted to do?" | A policy document. Ask instead for the actual scope on the credential, read from the system. The two are rarely the same, and the gap is the finding. |
| "How do you know it still works as intended?" | "It was tested before go-live." Eight months and a dozen silent model updates ago. |
| "Which AI systems are in scope?" | Three approved ones. Meanwhile staff are using twenty. Whether that is a scope decision or a nonconformity is genuinely unsettled, and it determines whether the certificate means anything. |
For U.S. federal environments: where FedRAMP fits
The Federal Risk and Authorization Management Program (FedRAMP) provides reusable security-assessment evidence for cloud service offerings used with federal information. It is important, but it answers a narrower question than this guide: whether a specific cloud service offering has completed the applicable federal assessment process. It does not certify an agency's complete AI-agent workflow, and it is not an AI-safety or agent-control framework.
The distinction matters because an agent extends beyond the model endpoint. Its authority path may include identity services, retrieval sources, memory, orchestration, tools, APIs, plug-ins, logging systems and downstream applications. A provider may have a FedRAMP-certified offering without every commercial feature, deployment, connector or external dependency being part of that same assessed offering. Verify the exact offering and use case rather than inferring coverage from the vendor name.
FedRAMP evidence can be inherited. Responsibility for the agency's actual system and use cannot.
The agency authorizing official still accepts risk for the federal information processed, the configuration selected, the integrations enabled, the agency-operated controls and the wider information system that uses the cloud service.
For an AI-agent deployment, defenders and authorizing teams should be able to answer five questions with evidence:
- Offering: Is the exact cloud service offering and certification status present in the FedRAMP Marketplace?
- Scope: What federal information is created, collected, processed, stored, transmitted or accessed across the workflow?
- Boundary: Which identity, orchestration, memory, tool, logging and downstream services handle that information or directly affect its confidentiality, integrity or availability?
- Responsibility: Which controls are inherited from the provider, and which remain with the agency or another service provider?
- Authorization evidence: Does the agency's authorization and continuous monitoring cover the complete authority path, or stop at the model or platform boundary?
This is where FedRAMP and the operating model in this book meet. FedRAMP supplies reusable assurance about a cloud offering. The agency still has to authorize and defend the system assembled around it.
Say this on Monday
"If an auditor asked us today why our agent did something in June, what would we hand them?"
Sourceslast verified 16 September 2026
- ISO/IEC 42001:2023, Information technology - Artificial intelligence - Management system. Clause and Annex A references throughout this part are to the 2023 edition. iso.org/standard/42001
- The standard distinguishes roles including the organization that develops AI and the organization that merely uses it; control references here are read against that distinction.
- FedRAMP, Using a FedRAMP Certified Cloud Service: certification applies to the cloud service offering; the agency authorizing official accepts risk for the agency's information, configuration, integrations and agency-operated controls. fedramp.gov/2026/agencies/use
- FedRAMP, Scope of FedRAMP: applicability depends on the agency use case and whether the service creates, collects, processes, stores or maintains federal information. fedramp.gov/2026/scope
- The FedRAMP Marketplace is the authoritative catalog for checking a cloud service offering's current certification status. fedramp.gov/marketplace
Turn the guidance into an AIMS roadmap
This chapter explains the responsibilities and evidence expected of organizations that use third-party AI. The companion AI governance page connects those operational questions with management-system scope, controls, implementation and assurance.
About the author
Jessen Kurien is a cybersecurity leader and the author of The Defender’s Guide to AI Agents. His 18+ years in cybersecurity include nearly 15 years at Microsoft, work as part of the founding team of the Microsoft Threat Intelligence Center, and detection engineering leadership in Microsoft Defender XDR. His work connects investigations, detection engineering and security operations with the evidence and accountability needed for AI security and governance.
This guide will go out of date.
Providers change how their logs work, models get retired, and new cases get disclosed. Ask to be told when this changes — no newsletter, just the updates.
Download the complete guide (PDF)
The telemetry contract, detection specifications, framework mappings and checklists are also published as files — the defender pack, CC BY 4.0, free to reuse.