Agent guardrails are not one layer. They are six authorities, and no one should hold them all
Identity, model access, containment, sequence, change approval and spend answer different questions. This week's releases show them separating. A framework for deciding who holds each one.
Listen to this article · 6 min
AI-generated narration of the full article.

Ask an enterprise team how its AI agents are governed and the answer is often one word: guardrails. Ask what a guardrail is, and the answers diverge. For one person it is a filter on the model’s output. For another it is a line in the system prompt. For a third it is the platform’s permission screen.
Meanwhile, in a single week, vendors shipped controls for agents from at least six different directions, and almost none of them is a guardrail in any of those senses. Read together, they suggest that agent governance is not a layer. It is a set of separate authorities, each answering a different question, and the most important design decision is who holds each one.
One week, six authorities
Identity: who is this agent, and who answers for it? RSA announced Agent ID on October 7 to discover agents, register them with named owners, risk tiers and lifecycle states, enforce policy at AI and MCP gateways with out-of-band approval, certify their access continuously and revoke it on decommissioning. The discovery and security module is due for general availability on November 16; the governance module in the first half of 2027.
Model access: which capabilities may this organization use, and for what? On October 6, Anthropic split its Cyber Verification Program into Defense, Red Team and Specialized tiers. Red Team access is for organizations only and takes weeks to review; Specialized access is reviewed with the US government. Access to a capability now depends on verified purpose, not only on holding an API key, an idea Anthropic first applied to life sciences.
Containment: what can this agent ever touch? Microsoft made Execution Containers generally available, so Windows enforces which files, networks and processes an agent reaches.
Sequence: may it do this now, given what it just did? AWS’s Strands Box makes authorization temporal: Dogwood policies read the agent’s history before allowing the next action.
Change approval: should this reach production? Cornerstone OnDemand’s database agents diagnose on their own but wait for a person before any destructive operation, with a five-minute window that defaults to deny.
Spend: how much may it consume before someone decides? GitLab added AI usage caps at subscription, group and user level, next to cost reporting by team, task and model.
| The question it answers | Natural holder | If the agent's builder holds it | |
|---|---|---|---|
| Identity | Who is it, and who answers for it? | Identity and access management | Nobody can revoke it cleanly |
| Model access | Which capabilities, for which purpose? | Risk or security leadership, with the provider | Capability follows whoever has a key |
| Containment | What can it ever touch? | Endpoint or platform security | A persuaded agent reaches what its host reaches |
| Sequence | May it act now, given what it just did? | Security engineering | Allowed steps add up to a forbidden outcome |
| Change approval | Should this reach production? | The owner who already approves changes | The author approves its own work |
| Spend | How much before someone decides? | Finance with the product owner | The invoice becomes the incident report |
Why they have to stay separate
They fail independently, or they fail together. A prompt injection defeats anything the agent reads and interprets. If containment and sequence rules live in the prompt, one malicious document disables both. Authorities enforced outside the agent, by the operating system, a gateway or a person, survive the agent being wrong. That is the argument behind our case that prompt guardrails are not a security boundary, extended to every authority.
They run on different clocks. Identity changes when someone leaves or a team reorganizes. Model access changes when a provider reclassifies a capability. Containment changes with each release of the agent. Approval happens per change, spend per month. One team cannot maintain six cadences well, and one platform setting cannot represent them.
Auditors already think this way. Segregation of duties is one of the oldest controls in finance and IT: the person who initiates a payment does not approve it. Agent governance is the same principle applied to a new kind of actor. An agent whose builder also defines its identity, its containment, its approvals and its budget is a segregation of duties finding waiting to be written.
Using it in a review
For each production agent, ask four questions.
- Who holds each of the six authorities? Write a name or a team, not a product.
- Do any two critical authorities rest with the team that builds the agent? Containment and change approval are the two that most often do.
- Can each authority act without the others? Revoke the identity without redeploying the agent. Deny an action without editing the prompt. Stop spending without deleting anything.
- What happens when an authority is silent? An approval that times out, a gateway that cannot reach its policy service, a budget service that is down. Default deny should be a decision, not an accident.
This refines a view we have held since our argument that agents need an ID, a manager, a budget and a file. The file records the agent. The authorities enforce it, and they should not all answer to the same person. It also gives shape to autonomy that is earned: promotion up the ladder is a decision these authorities make together, not a setting one of them changes alone.
Our expectation, and it is a view rather than a fact, is that buyers will start asking vendors a simple question: which of the six does your product hold, and which does it leave to us? Products sold as “the guardrails” will have to answer it.