The Defender’s Guide to AI Agents·Chapter 8
When the Model Changes Underneath You
By Jessen Kurien · CISM, CISA, ISO/IEC 42001 Lead Implementer & Lead Auditor ·
The moving target
Everything so far assumed the agent stays the same. It does not, and this is the part of the problem that has no equivalent anywhere else in security.
You are used to software that changes when you change it. This changes when someone else changes it, on their schedule, and often the name you have written in your code stays exactly the same while the thing behind it does not.
The best-documented case is also the most instructive, because it is not the story people tell about it. Researchers compared two dated editions of the same model family three months apart and found the share of code it produced that would run as written fell from 52% to 10%. The popular version of that finding is “the model got worse.” The authors’ own explanation is more useful to you: the later edition had started wrapping its answers in markdown formatting, which broke everything downstream that consumed the output directly. The reasoning was fine. The shape of the answer changed, and that was enough.
Sit with that, because it is the failure mode to plan for. It would pass any eyeball review - a human reading those answers would call them good. It only breaks the things reading them automatically, which is most of what an agent does.
Four things to do about it. They are not difficult; they are just unfamiliar.
Pin the version
Model names come in two kinds, and the difference is everything. A pinned name points at one specific frozen edition, usually with a date in it. A floating name is a convenient shorthand that quietly resolves to whatever the newest edition happens to be today.
Never send a floating name to production. If you do, the thing that answered you was chosen at the moment of asking, by someone else.
Providers differ enough that you have to check yours specifically, and the traps are not where you expect:
| Where it runs | How pinning works | The trap |
|---|---|---|
| Direct from a model provider | Dated editions alongside floating shorthand names. Some providers have moved to making every name a pinned one. | Old code still carrying a floating name that nobody has looked at in a year. |
| Through a cloud platform | Usually a per-deployment setting: never upgrade, upgrade when a new default appears, or upgrade only at retirement. | The default is often to upgrade automatically. Choosing "never" is right, but it means the deployment stops working at retirement rather than drifting - so the retirement date has to be in someone's calendar. |
| Through a model marketplace | Models move through states: current, on notice, withdrawn. The state is readable through the platform's own interface. | On at least one major platform, once a model is on notice, an organization that has not used it for a couple of weeks can lose access to it entirely. A fallback you never call is a fallback you do not have. |
Which leads to a rule that sounds odd and is not: exercise your fallback on a schedule. Once a week, send one real request to the previous pinned version. It proves the fallback still works, and on some platforms it is the only thing keeping it available.
Record what was running
When something goes wrong in March you will want to know what the agent was in January. The model version alone will not tell you, because two other things change more often than the model does - and both of them are yours.
What was running is three things, not one.
The model edition that actually answered - read from the response, not from what you asked for.
The standing instructions - the system prompt, which your own team edits far more often than the provider ships a model.
The list of tools that were available - which changes whenever anyone adds a plug-in.
Take a fingerprint of all three together and attach it to every consequential action. Then "when did this start" becomes a search rather than an investigation. Log only the model and you will find the same model on both the good days and the bad ones, and conclude that nothing changed.
Keep a set of tests and re-run it
Not a benchmark. You are not measuring how clever the model is. You are watching for the day it becomes a different model, which is exactly what a synthetic transaction does for a payment system.
- Thirty to eighty cases per agent is enough. You are looking for a collapse, not a nuance. A fall of the size described above is unmissable in forty cases.
- Build them from real traffic. Take sessions that went right, freeze them with the outcome you expected. Cases you invent test a world you imagined.
- Cover four things. Does it still do the job. Does it still refuse what it refused last month, including your collection of injection attempts. Does it still pick the right tool - this degrades first and is invisible in the quality of the writing. Does its structured output still parse.
- Mark answers automatically wherever you can. Tool name matched or it did not; the output parsed or it did not. No judgment needed.
- If you use a model to grade, pin that model too, separately. A drifting examiner produces drifting marks and you will chase the wrong thing for a week.
- Run it on a schedule, not only on change. This is the whole point. A test that runs when you change something cannot catch a change you were never told about.
File every result against the fingerprint from Chapter 8. That eval set is a security control, not a quality one - silent decline is how these deployments actually fail, long before anyone attacks them.
Treat a change as a change
Three kinds, each needing a different path, and most organizations have a path for only the first.
Class A · you changed it
A new model version, a rewritten prompt, a new tool. A routine change with a pre-approved test plan - and the test plan is the eval set. Passes, ships. This should need no meeting.
Class B · they announced it
A retirement date, a new default, a deprecation. A normal change with a real date on it. Watch the provider's deprecation pages, their service-health alerts, and poll the platform weekly for models moving onto notice - there is usually no push notification.
Class C · it changed silently
Found by the mismatch alert from Chapter 6, or by the scheduled eval dropping. This is an unplanned change to a production system, which is an incident, not a change. Almost nobody has this path, which is why silent decline runs for weeks.
And the part that outlives any tooling: put notice periods in the contract. The spread across providers is stark - some commit to months, some to a fortnight, and at least one has historically given a median of zero days. That is a supplier conversation, and it is the one control here a security team cannot implement alone.
Say this on Monday
"Can we prove which version of the model answered, on any given day last month? If not, we can't investigate anything that happened last month."
SourcesLast verified 17 September 2026
- Behavioral drift between two dated editions of the same model family, and the authors’ attribution of the executability drop to output formatting: Chen, Zaharia and Zou, “How Is ChatGPT’s Behavior Changing over Time?”, arXiv, 2023. arxiv.org/abs/2307.09009
- Alias versus pinned model identifiers, and the commitment that an existing model ID is not updated. platform.claude.com
- Per-deployment upgrade policies, and the notice given before a default version changes. learn.microsoft.com
- Model lifecycle states, notice periods, and the loss of access after a period of inactivity on a legacy model. docs.aws.amazon.com
About the author
Jessen Kurien is a cybersecurity leader and the author of The Defender’s Guide to AI Agents. His 18+ years in cybersecurity include nearly 15 years at Microsoft, work as part of the founding team of the Microsoft Threat Intelligence Center, and detection engineering leadership in Microsoft Defender XDR. His work connects investigations, detection engineering and security operations with the evidence and accountability needed for AI security and governance.
This guide will go out of date.
Providers change how their logs work, models get retired, and new cases get disclosed. Ask to be told when this changes — no newsletter, just the updates.
Download the complete guide (PDF)
The telemetry contract, detection specifications, framework mappings and checklists are also published as files — the defender pack, CC BY 4.0, free to reuse.