← All insights

News analysisLens: United States5 min read

When an AI agent becomes your SRE, permissions become reliability engineering

AWS DevOps Agent investigates incidents across AWS, other clouds and on premises, and the new Well-Architected Agent reviews architecture in preview. Both show where operational agents draw the line.

Listen to this article · 8 min

AI-generated narration of the full article.

An old brass pressure gauge on a pipe, in La Madre duotone, beside the words Agents in operations
Photo: rawpixel (CC0)

A chat assistant that gives a wrong answer wastes a few minutes. An operations agent that takes a wrong action during an incident can turn a partial outage into a full one. That is why the most important design decisions for operational agents are not about the model. They are about what the agent may read, what it may propose and what it may do.

AWS has two agents that make this concrete. AWS DevOps Agent, generally available since March 2026 and with its API reference updated on October 2, works as an always-available operations teammate. AWS Well-Architected Agent, announced in preview on October 1, brings architecture review into the same pattern.

What each agent does

AWS describes DevOps Agent as an agent that resolves and proactively prevents incidents, optimizes reliability and performance, and handles on-demand SRE tasks across AWS, multicloud and on-premises environments. It works with observability tools, runbooks, code repositories and CI/CD pipelines, and correlates telemetry, code and deployment data. Integrations include GitHub and GitLab, Dynatrace, Datadog, New Relic and Splunk, plus MCP servers.

Well-Architected Agent is AWS’s AI evolution of Trusted Advisor and the Well-Architected Tool. It reads resource configurations, utilization metrics and application topology through customer-managed IAM roles, analyzes them against Well-Architected practices across more than 65 services, and prioritizes findings against the business goals the customer declares. It can also review Terraform, CloudFormation and CDK projects before deployment. Findings come with an implementation package: SSM runbooks, CLI scripts or guided console steps. The preview runs in US East (N. Virginia), US East (Ohio) and US West (Oregon), can onboard workloads from any commercial region, and requires an AWS Support plan.

The boundaries AWS chose

Read the DevOps Agent changelog for the last few months and a clear permission model emerges:

  • Read is the default. Read-only actions are available out of the box. Directed actions that create or modify resources are disabled by default and must be enabled explicitly, through a managed policy and the Agent Space preferences.
  • Proposals before actions. When an alarm triggers an investigation, the agent presents mitigation proposals that an operator can review and refine before applying.
  • Every approval is attributable. AWS says each approval and action is attributable to the approving operator in AWS CloudTrail.
  • Code runs in a box. Investigation code runs in an isolated microVM per investigation, with AWS calls limited to read-only operations.
  • The customer picks write access to code. A custom GitHub App can be created with read-only or read-and-write access, owned by the customer.

Well-Architected Agent goes further in the conservative direction: it makes no changes on its own. The customer chooses the remediation path and executes it.

Operational agent: where authority sitsON BY DEFAULTOPT-IN AND ACCOUNTABLE01Readtelemetry,code,deploys02Investigateandcorrelate03Proposemitigationor fix04Operatorapproves05Directedaction, ifenabled06AttributedinCloudTrailAgent autonomyHuman authority and evidence
  1. Read telemetry, code, deploys
  2. Investigate and correlate
  3. Propose mitigation or fix
  4. Operator approves
  5. Directed action, if enabled
  6. Attributed in CloudTrail
  • On by default: Read telemetry, code, deploys · Investigate and correlate · Propose mitigation or fix
  • Opt-in and accountable: Operator approves · Directed action, if enabled · Attributed in CloudTrail

Agent autonomyHuman authority and evidence

The agent does the legwork by default. Changes to production need an explicit grant and leave a named approver.

Why permissions are now a reliability decision

In classic SRE practice, reliability comes from limiting blast radius: small changes, staged rollouts, quick rollback. An operational agent adds a new source of change, so the same discipline applies to it:

Scope the agent like a production service account, not like a senior engineer. An Agent Space defines which accounts and resources the agent can see. Start narrow. The pattern matches the bounded autonomy we saw in the agentic SOC: autonomy on reading, authority on writing.

Give the agent its own identity. If agent actions run under a shared admin role, nobody can tell afterward what the agent did and what a person did. This is the case for treating agent identity as its own IAM discipline.

Treat agent-applied mitigations as changes. For a U.S. public company, a change to systems that support financial reporting falls under SOX IT general controls: it needs authorization and evidence. A CloudTrail record naming the approving operator helps; the change ticket and the post-incident review still need to exist.

Watch the cross-boundary reach. An agent that reads Azure resources, on-premises systems and third-party observability tools holds credentials into all of them. Those credentials need the same rotation, scoping and review as any integration.

Know where the data goes. Agent Space data (investigations, topology, recommendations) is stored in the region where it is created, and inference for U.S. Agent Spaces stays within the United States. AWS’s security documentation adds two details worth a line in your risk register: region-restricting service control policies do not apply to the agent’s inference, and the agent does not filter personal data when it summarizes logs, so redaction has to happen before data reaches your observability stack.

What changes for architecture review

Well-Architected reviews have usually been periodic questionnaires. An agent that reads live configuration and IaC turns them into something closer to continuous review: drift becomes visible when it happens, not at the next quarterly session. The accountability does not move, though. The agent recommends and packages the fix; the architect decides, and the change goes through the pipeline like any other, ideally with the IaC review step catching problems before deployment, the same way Google’s agents work inside code review.

What to do now

  1. Write the permission ladder before turning anything on: read, propose, act with approval, act alone. Assign each action type to a rung.
  2. Keep directed actions off until you have a playbook that says which mitigations the agent may execute and who approves.
  3. Create a dedicated role per Agent Space and scope it to the accounts it needs.
  4. Route agent proposals into your change process so an approved mitigation carries a ticket, not only a CloudTrail line.
  5. Pilot the Well-Architected Agent on one workload with declared goals, and compare its findings with your last manual review.
  6. Measure the assist: time to first useful hypothesis, share of proposals accepted, proposals that would have made things worse.

The bottom line

AWS’s operational agents show a sensible default: read freely, propose generously, act only with an explicit grant and a named approver. Enterprises should adopt that default as policy, not just as a product setting, because the next operational agent they deploy may not ship with the same restraint.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case