← All insights

Architecture guideLens: United States4 min read

Agent autonomy should be earned, not configured. A ladder for granting it one action at a time

Autonomy is usually set once, per agent, in a settings screen. La Madre's view: grant it per action class, on evidence, and take it back when the model or the risk changes.

Listen to this article · 6 min

AI-generated narration of the full article.

A metal ladder descending into calm sea water at dusk, in La Madre duotone, beside the words Autonomy is earned
Photo: Samuel Zeller (StockSnap, CC0)

In most agent platforms, autonomy is a setting. You choose whether the agent asks before acting, tick a box, and the decision is made for every action the agent will ever take. It is easy to configure and hard to defend, because nothing about the agent’s track record went into it.

This week Gartner put the problem plainly. In an article published October 2, it argued that the push for autonomous agents is moving faster than enterprise readiness, and repeated its prediction that 40% of agentic AI projects will be canceled by the end of 2027 because of escalating costs, unclear business value or inadequate risk controls. “The challenge is not creating more capable AI agents,” said Erick Brethenoux, a Distinguished VP Analyst at Gartner. “It’s creating agents that are reliable enough to earn greater autonomy while operating within clear accountability and governance boundaries.”

We agree, and we want to make it operational. Here is La Madre’s view of how autonomy should be granted, with the caveat that this is a framework built from what we have seen in the market, not a standard anyone has adopted.

Three signals that point the same way

Our position comes from a pattern across recent coverage:

  • Security operations got there first. In the agentic SOC, agents gather context and propose; analysts decide and contain.
  • Operational agents ship with action off. AWS DevOps Agent has read-only actions on by default and actions that create or modify resources off by default, with each approval attributed to a named operator.
  • Model providers now tell builders to set decision boundaries. OpenAI’s GPT-6 guide recommends stating which actions can proceed independently and which need approval, replacing blanket “always ask” rules.

And the counterexample: when an agent keeps acting after the person who started it leaves, as with Copilot Autopilot, the ownership question arrives before the evidence does.

The ladder

Five rungs, granted per action class011. Read andsummarize022. Recommend033. Act withapproval eachtime044. Act withinbounds,sampled review055. Actbroadly,monitoredNo change to the worldHuman approves each action
  1. 1. Read and summarize
  2. 2. Recommend
  3. 3. Act with approval each time
  4. 4. Act within bounds, sampled review
  5. 5. Act broadly, monitored

No change to the worldHuman approves each action

An agent can sit on rung 4 for one action and rung 2 for another. Evidence moves it up; a model change or an incident moves it down.
  1. Read and summarize. The agent gathers and explains. Nothing changes in any system.
  2. Recommend. The agent proposes a specific action with its reasoning and evidence. A person acts.
  3. Act with approval. The agent executes, but only after a named person approves each action.
  4. Act within bounds. The agent executes pre-approved action classes alone, inside limits (amount, scope, volume, reversibility), with a sample of actions reviewed after the fact.
  5. Act broadly. The agent decides and executes across a wider set of actions, monitored continuously. Few enterprise workflows should be here in 2026.

Four rules that make the ladder work

Grant autonomy per action class, not per agent. “Reset a password” and “close a customer account” are different risks even when the same agent does both. The unit of trust is the action class.

Promote on evidence. Moving an action class from rung 3 to rung 4 should require a record: how many approvals, how many rejections, how many corrections, over what period. If approvers accept 99% of proposals without edits over a few hundred cases, that is evidence. If they rubber-stamp without reading, it is not, which is why approval quality needs its own observability.

Demote automatically. Three events should send an action class back a rung: an incident, a material change in what the action touches, and a change of model. The evidence was earned by one model; a replacement has not earned it yet. This is the link we see between autonomy and the model lifecycle problem we described in our piece on model retirement.

Write it in the agent’s file. We proposed that every production agent carry a file with identity, sponsor, route, budget and record of actions. The autonomy rung for each action class belongs in it, with the date and the evidence behind it.

What decides the ceiling

Not every action class should climb. Four questions set the ceiling before any evidence is collected:

  • Is it reversible? Refunding a fee can be undone; sending a regulatory filing cannot.
  • What is the blast radius? One record, one customer, or every customer?
  • Can we see it quickly? Autonomy without monitoring is just hope.
  • Who answers for it? If no named person would accept accountability for the agent’s action, the ceiling is rung 3.

In the United States there is no federal rule that tells an enterprise how much autonomy an agent may have. The ceiling is set by existing obligations (financial controls, consumer protection, sector regulators) and by the organization’s own risk appetite. That makes writing it down more important, not less, because the auditor’s first question will be how the decision was made.

What to do now

  1. List the action classes each production agent can perform, not just the agents.
  2. Assign a rung and a ceiling to each, using the four ceiling questions.
  3. Start counting: approvals, rejections, edits and incidents per action class.
  4. Set promotion criteria in advance, so moving up is a decision with evidence, not a mood.
  5. Wire demotion to model changes, so a replacement model starts one rung lower until it proves itself.

The bottom line

Gartner’s warning is about pace: autonomy is being configured faster than reliability is being proven. The fix is not to slow everything down. It is to make autonomy something an agent earns, one action class at a time, on evidence a skeptic would accept, and loses when the facts change.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case