← All insights

Case analysisLens: United States4 min read

Google's most useful AI agents don't chat. They sit inside code review and block vulnerabilities

Google describes scanning, triage and fix agents embedded in its development lifecycle, with deterministic checks and human review. What enterprises can copy from agents built into existing controls.

Listen to this article · 6 min

AI-generated narration of the full article.

A rusty padlock closing an old green door, in La Madre duotone, beside the words Agents as controls
Photo: Michal Jarmoluk (StockSnap, CC0)

When enterprises picture an AI agent, they usually picture a conversation: someone asks, the agent answers or acts. Some of the most valuable agents look nothing like that. They have no chat window. They run inside a process the company already trusts, and their output is a decision that process already knows how to handle.

A September post by two Google engineering leaders, a Distinguished Engineer and an Engineering Fellow, describes exactly that: AI agents embedded in the software development lifecycle that secures Google’s own infrastructure code. It is one of the more detailed public accounts of agents operating inside a control process at scale, and its design choices transfer well beyond Google.

How the system works

According to Google, the system has three stages and a fix step:

  • Pre-submit scanning. Every code check-in, across every layer of the stack, is evaluated in real time by AI agents inside developer tools. The agents use localized threat models built with live codebase metadata and dependency call graphs.
  • Triage. A separate, specialized triage agent validates each finding in a lightweight scan of under a minute, using abstract syntax tree parsing, call-graph traversal and pre-indexed safety rules to prove that the vulnerable path is actually reachable by an attacker. Google reports over 92% precision at this stage.
  • Post-submit scanning. A nightly scan, running on off-peak capacity, looks for vulnerabilities introduced across multiple changes.
  • Automated fixes, human approval. A bug-fix agent uses the findings and generated proofs to construct a fix, which is submitted for human review as part of the original change request’s review.

Google says the system covers every code change deployed onto its infrastructure, across hundreds of millions of lines of code, and prevents hundreds of vulnerabilities per month from reaching the codebase or production. It reports false positives as low as 3% in some cases. These are Google’s own figures about its own environment. The work builds on Mantis, a multi-agent review harness Google has released as open source.

Agents inside an existing code review process01Developersubmits achange02Scanningagent withthreatmodel03Triageagent withastructuralproof04Fix agentproposes apatch05Reviewerapproves inthe samereview06Nightlyscan acrosschangesAI agent stepHuman decision, in the process that already existed
  1. Developer submits a change
  2. Scanning agent with threat model
  3. Triage agent with a structural proof
  4. Fix agent proposes a patch
  5. Reviewer approves in the same review
  6. Nightly scan across changes

AI agent stepHuman decision, in the process that already existed

No new interface, no new approval queue. The agents feed a control the organization already trusts.

Five design choices worth copying

1. Separate the agents. Google’s first recommendation is to keep the harnesses, rules and context of development, scanning and triage agents separate, to prevent bias. It is the agent version of separation of duties: the agent that finds a problem should not be the one that decides it is real, and neither should write the fix unchecked.

2. Pair AI with deterministic checks. The triage step does not trust a model’s opinion that code is vulnerable. It uses structural analysis to prove reachability. Combining a fast AI scan with deterministic validation is what drives down both latency and false positives. That pattern applies to any agent whose output triggers work for people.

3. Feed agents the context you already have. Google points to existing threat models as the key input for precision. Most enterprises have equivalent assets: control catalogs, architecture decision records, data classification. They are often the best context an agent can get.

4. Put human approval where humans already decide. The fix arrives in the original change review. Reviewers do not learn a new tool or work a new queue. Adoption follows because the agent fits the workflow, not the other way around.

5. Use the harness to absorb model variability. Google notes that a multi-agent harness can compensate for variability in model choice. The design, not a specific model, carries the reliability. We made a similar argument in our analysis of reusable AI harnesses.

What an enterprise can realistically do

Few enterprises have Google’s scale or its security engineering bench. The lesson is not to rebuild this system. It is to look for controls you already run where an agent could do the first pass:

  1. Pick one existing control with a clear decision: code review, access review, vendor risk questionnaires, change approval.
  2. Split finding from confirming. Use one agent to flag, a deterministic check or a second agent to confirm, and measure precision before anyone sees the output.
  3. Deliver results inside the existing workflow, with the same human approver and the same evidence trail.
  4. Track false positives as a first-class metric. An agent that floods reviewers will be switched off, regardless of how many real issues it finds.
  5. Keep the agents’ own changes under review. As we noted in our analysis of Palantir’s AIP Evolve, agents that change systems need evaluation and human gates of their own.

The bottom line

The agents in Google’s account are invisible to most of the people they help. That is the point. High-value enterprise agents may increasingly disappear into existing control processes, where success is measured by precision, time to fix and reviewer trust rather than by conversations.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case