Cornerstone cut database diagnosis from 45 to 10 minutes. Its agents never touch production alone
Cornerstone OnDemand's database agents gather evidence across SQL Server, Jira and dashboards, while destructive actions wait for a person and default to deny. An AWS pattern formalizes the split.
Listen to this article · 5 min
AI-generated narration of the full article.

A query locks a table. A second query waits on it, a third waits on the second, and somewhere a page loads slowly for customers. The engineer on call opens system views, logs and dashboards, looking for the head of the chain. At Cornerstone OnDemand, that hunt used to take about 45 minutes per incident.
An AWS case study published on October 7 describes what the company built to shorten it, and the most useful detail is not the speed. It is where the agents stop.
Engineers as the integration layer
Before the project, diagnosis was manual: querying system views, cross-referencing logs, coordinating between teams. Database lifecycle workflows took ten or more manual steps. Reports between site reliability and data teams lagged by about fifteen minutes, and redundant alerts competed for attention.
None of that is a model problem. It is an evidence assembly problem, spread across systems that do not talk to each other, with a person acting as the glue.
Thirteen narrow agents and one orchestrator
Cornerstone’s system, called Orion AI, uses an orchestrator that routes each request to one of thirteen domain agents, among them database diagnostics, session blocking analysis and real-time SQL diagnostics. It runs on Amazon Bedrock with the open source Strands Agents framework, and the agents share tools through MCP. They read SQL Server through an internal operations API, and they work with Jira, metrics dashboards and on-call schedules.
The agents map blocking chains to root causes, open Jira tickets already filled in and assigned, and correlate and deduplicate alerts. According to figures from Cornerstone’s internal data operations team, published by AWS, average diagnosis time fell from about 45 minutes to about 10, a 78% reduction, and redundant alerts fell by a median of 65%.
Diagnosis time is not resolution time. The case reports one segment of the incident, the one the agents were built for.
Where Cornerstone's agents act and where they ask
Agents act alone
- Query system views
- Map blocking chains
- Correlate and deduplicate alerts
- Open and assign Jira tickets
Read and record
A person decides
- Destructive database operations
- Anything a guardrail flags as critical
Five-minute window, denied on timeout
The controls are the design
Three choices in the case are worth copying, and none of them depends on Bedrock.
Approval has a clock, and silence means no. Destructive operations pause until a person confirms, within a five-minute window, and the default on timeout is deny. An approval with no deadline becomes a queue. An approval that defaults to allow is not an approval.
Memory is not trusted for live data. The team scoped conversational memory to the session and bypassed it entirely for live metrics, so an agent never answers a question about the database’s current state from what it remembered ten minutes ago.
Agents are split by domain, not by difficulty. Each agent gets a narrow set of tools, which the team says improved tool selection. Routing defaults to keyword matching, which handles about 80% of requests, with semantic search as the fallback.
It is also a modest project, which makes it more credible. A three-person team delivered it in six months, on Amazon ECS because the Bedrock AgentCore runtime was not available when they started.
The same split, as a reference pattern
On the same day, AWS published an architecture that turns Cornerstone’s instinct into a template. When an AWS DevOps Agent investigation finishes, an event starts a Lambda durable function, and a Bedrock model proposes a remediation. Read-only tools run on their own. Infrastructure changes suspend the workflow until a person approves. The model can only invoke tools on a curated allowlist, approvals are checkpointed, and the execution history shows each step, each tool call and the model’s reasoning.
The post does not say what happens when a person rejects the fix. That is the part each team has to design, and Cornerstone’s default deny is a reasonable answer.
We wrote earlier that when an agent becomes your SRE, permissions become reliability engineering. These two pieces show what that looks like in practice: investigation sits on a high rung of the autonomy ladder, and every change sits one rung lower, behind a person and a clock. The lesson for anyone building an operations agent is to measure it by how fast it puts the right evidence in front of a person, and to let it earn more only on that evidence.