Google's Agent Substrate is built for fleets of agents, not one. That changes what platform teams own
GKE Agent Substrate suspends idle agents to storage and resumes them in isolated sandboxes in under a second. Why agent execution is becoming a platform problem, and what it leaves open.
Listen to this article · 6 min
AI-generated narration of the full article.

Most enterprise agent work so far has been about building one agent well: the prompt, the tools, the evaluation. The next problem is less glamorous. What happens when a company runs thousands of agents, each one executing code, most of them idle most of the time, and every one of them needing to be isolated from the rest of the infrastructure?
Google’s answer for Kubernetes is GKE Agent Substrate, highlighted in its September AI infrastructure recap. It is a useful signal of where the bottleneck is moving: from model intelligence to safe, economical execution.
What Agent Substrate does
Google describes Agent Substrate as an open-source execution runtime for agentic workloads on GKE. Its design starts from how agents actually behave. Google’s earlier announcement put it plainly: agents run in short bursts, then sit idle for long periods, and they need strong kernel and network isolation because they execute code that nobody reviewed line by line.
Instead of keeping one container alive per agent, Agent Substrate decouples agent state from the underlying pods. Suspended agents are stored as snapshots, and resumed on demand onto a shared pool of warm workers. Working memory and files are preserved, so an agent resumes where it paused.
Google’s claims for it are aggressive: it says Agent Substrate can run millions of sandboxes at 10 times the density of standard container runtimes, with resume operations under 500 milliseconds at more than 500 suspend and resume activations per second. Isolation uses gVisor, with Cloud Hypervisor as an option.
- Request for an agent
- Snapshot restored onto a warm worker
- Work runs in an isolated sandbox
- Agent suspended to a snapshot
- Worker returns to the pool
Isolation boundary for AI-generated codeAgent memory and files at rest: data to classify and protect
Read the status carefully
The documentation is more precise than the headline. Agent Substrate is available to all Google Cloud customers for evaluation and non-production use; production support is offered on an allowlist basis under a private GA program. It runs only on GKE Standard clusters, not Autopilot, and requires Workload Identity Federation.
It also lists limitations that matter for enterprise designs: no GPU passthrough, open network connections are not preserved across suspend and resume, and egress policies are not supported. Google says it is offered at no extra charge in GKE.
Its sibling, GKE Agent Sandbox, is generally available, and Google added a version optimized for reinforcement learning in September.
What this means for platform teams
A managed agent runtime from a model provider, such as OpenAI’s Agents API, hides execution behind an API. Agent Substrate is the opposite choice: execution stays in your clusters, and your platform team owns it. That brings control, and a list of decisions.
Isolation is the baseline, not the design. A gVisor sandbox protects the host from AI-generated code. It does not decide what that code may reach. With egress policies unsupported today, network controls have to come from elsewhere in your architecture before anything sensitive runs there.
Snapshots are data. If an agent’s memory and files are preserved between runs, the snapshot store holds whatever the agent was working on: customer records, credentials in environment variables, intermediate results. Classification, encryption, retention and access to that storage need an owner, just as a database would.
Identity per agent is still your job. The documentation requires Workload Identity Federation for the system, but does not describe a per-agent identity model. If thousands of agents share one service account, your audit trail cannot tell them apart.
Idle is the new cost driver. The economics of agent fleets depend on how much you pay for agents that are waiting. Suspend and resume is a cost lever, and capacity planning becomes a shared conversation between platform and FinOps teams.
For US enterprises
Teams already running large GKE estates can start learning now, with the right boundaries:
- Evaluate in a non-production project, as the documentation intends, with synthetic data.
- Design network egress controls first, since the runtime does not enforce egress policies.
- Treat the snapshot bucket as a sensitive data store: encryption keys, retention rules, access reviews.
- Decide the identity model before the number of agents grows: one identity per agent or per agent type, never one for all.
- Measure idle time in your current agent workloads to estimate what density could actually save.
The bottom line
Agent Substrate is infrastructure for a stage most enterprises have not reached yet: fleets of agents, not pilots. But the questions it raises apply now. Where does agent state live, what can agent code reach, and who answers for each agent? Those are the same controls we map in our 2026 enterprise AI stack guide. Platform teams that answer them early will be the ones able to scale.