Google's most useful AI agents don't chat. They sit inside code review and block vulnerabilities
Google describes scanning, triage and fix agents embedded in its development lifecycle, with deterministic checks and human review. What enterprises can copy from agents built into existing controls.
Listen to this article · 6 min
AI-generated narration of the full article.

When enterprises picture an AI agent, they usually picture a conversation: someone asks, the agent answers or acts. Some of the most valuable agents look nothing like that. They have no chat window. They run inside a process the company already trusts, and their output is a decision that process already knows how to handle.
A September post by two Google engineering leaders, a Distinguished Engineer and an Engineering Fellow, describes exactly that: AI agents embedded in the software development lifecycle that secures Google’s own infrastructure code. It is one of the more detailed public accounts of agents operating inside a control process at scale, and its design choices transfer well beyond Google.
How the system works
According to Google, the system has three stages and a fix step:
- Pre-submit scanning. Every code check-in, across every layer of the stack, is evaluated in real time by AI agents inside developer tools. The agents use localized threat models built with live codebase metadata and dependency call graphs.
- Triage. A separate, specialized triage agent validates each finding in a lightweight scan of under a minute, using abstract syntax tree parsing, call-graph traversal and pre-indexed safety rules to prove that the vulnerable path is actually reachable by an attacker. Google reports over 92% precision at this stage.
- Post-submit scanning. A nightly scan, running on off-peak capacity, looks for vulnerabilities introduced across multiple changes.
- Automated fixes, human approval. A bug-fix agent uses the findings and generated proofs to construct a fix, which is submitted for human review as part of the original change request’s review.
Google says the system covers every code change deployed onto its infrastructure, across hundreds of millions of lines of code, and prevents hundreds of vulnerabilities per month from reaching the codebase or production. It reports false positives as low as 3% in some cases. These are Google’s own figures about its own environment. The work builds on Mantis, a multi-agent review harness Google has released as open source.
- Developer submits a change
- Scanning agent with threat model
- Triage agent with a structural proof
- Fix agent proposes a patch
- Reviewer approves in the same review
- Nightly scan across changes
AI agent stepHuman decision, in the process that already existed
Five design choices worth copying
1. Separate the agents. Google’s first recommendation is to keep the harnesses, rules and context of development, scanning and triage agents separate, to prevent bias. It is the agent version of separation of duties: the agent that finds a problem should not be the one that decides it is real, and neither should write the fix unchecked.
2. Pair AI with deterministic checks. The triage step does not trust a model’s opinion that code is vulnerable. It uses structural analysis to prove reachability. Combining a fast AI scan with deterministic validation is what drives down both latency and false positives. That pattern applies to any agent whose output triggers work for people.
3. Feed agents the context you already have. Google points to existing threat models as the key input for precision. Most enterprises have equivalent assets: control catalogs, architecture decision records, data classification. They are often the best context an agent can get.
4. Put human approval where humans already decide. The fix arrives in the original change review. Reviewers do not learn a new tool or work a new queue. Adoption follows because the agent fits the workflow, not the other way around.
5. Use the harness to absorb model variability. Google notes that a multi-agent harness can compensate for variability in model choice. The design, not a specific model, carries the reliability. We made a similar argument in our analysis of reusable AI harnesses.
What an enterprise can realistically do
Few enterprises have Google’s scale or its security engineering bench. The lesson is not to rebuild this system. It is to look for controls you already run where an agent could do the first pass:
- Pick one existing control with a clear decision: code review, access review, vendor risk questionnaires, change approval.
- Split finding from confirming. Use one agent to flag, a deterministic check or a second agent to confirm, and measure precision before anyone sees the output.
- Deliver results inside the existing workflow, with the same human approver and the same evidence trail.
- Track false positives as a first-class metric. An agent that floods reviewers will be switched off, regardless of how many real issues it finds.
- Keep the agents’ own changes under review. As we noted in our analysis of Palantir’s AIP Evolve, agents that change systems need evaluation and human gates of their own.
The bottom line
The agents in Google’s account are invisible to most of the people they help. That is the point. High-value enterprise agents may increasingly disappear into existing control processes, where success is measured by precision, time to fix and reviewer trust rather than by conversations.