← All insights

Architecture guideLens: United States4 min read

Your agent is slow, and it may not be the model. Agents put a new kind of load on the data path

Google now pitches storage and network instances for agents that query vector stores and operational databases. The point is broader: agent latency and reliability live in the data path.

Listen to this article · 5 min

AI-generated narration of the full article.

Light trails of traffic curving along a highway at night, in La Madre duotone, beside the words The data path
Photo: rawpixel (CC0)

When an agent feels slow, teams usually blame the model and start shopping for a faster one. Often the model is a third of the problem. The rest is the path between the agent and the data it needs: vector searches, database queries, API calls, each repeated several times per task. Google made the point in passing in its October 5 infrastructure announcement, and it deserves more attention than the instance specifications that came with it.

What Google said

In the Google Cloud Modernize post by Souvik Choudhury and Tom Nikl, Google wrote that “with the rise of real-time AI agents querying backend systems, workloads increasingly require high-throughput infrastructure that removes I/O bottlenecks.” It positioned two storage-heavy machine families for that traffic: Z4D, now GA, with up to 84,000 GiB of local NVMe SSD and 400 Gbps networking, and Z4M, in Preview, with up to 168,000 GiB of local NVMe, 400 Gbps networking and RDMA support. The stated goal is to reduce I/O wait and prevent query timeouts “when real-time AI agents query vector stores, operational databases, and large-scale data pipelines.”

You do not need these machines to take the lesson. Agents create a different access pattern from the applications your data platforms were built for.

Anatomy of one agent task

Where the time goes in a typical agent task01Model plansthe steps02Vectorsearch forcontext03Operationaldatabaselookup04Modelreasons onresults05API call toact orverify06Modelwrites theanswerModel timeData path time
  1. Model plans the steps
  2. Vector search for context
  3. Operational database lookup
  4. Model reasons on results
  5. API call to act or verify
  6. Model writes the answer

Model timeData path time

Half the steps are not model calls. Each one adds latency, and the slowest call in a chain sets the experience.

A person using an application makes a query, reads the result and decides what to do next. An agent can make dozens of reads to finish one task, often in parallel, and it does not pause to think between them. Three effects follow:

  • Latency compounds. Six sequential steps at 300 milliseconds each is nearly two seconds before any model time. A step that is slow 1% of the time makes the whole task slow far more often, because every task crosses many steps.
  • Concurrency multiplies. A thousand employees with an agent each can generate the query load of a much larger user base. Connection pools, read capacity and rate limits sized for human traffic run out first.
  • Timeouts become wrong answers. When a retrieval call times out, many agents do not fail; they continue with less context and answer anyway. A data path problem shows up as a quality problem.

Design for agent traffic

Trace every step. Instrument the agent so each model call, retrieval, query and API call appears as a span with its own latency. Without that, you will optimize the model while the database waits. We argued the same for quality in our piece on measuring retrieval as its own system.

Give agents their own lane. Route agent reads to replicas, caches or purpose-built read stores rather than the primary transactional database. Give agent identities their own rate limits and connection pools, so a busy agent fleet cannot slow the checkout page.

Keep the hops short. Put the vector index, the operational data the agent reads and the model endpoint in the same region where possible. Every cross-region hop is paid on every step of every task.

Budget latency per step. Set a target for each step type and alert on it, the way you would for a service-level objective. Decide what the agent does when a step misses its budget: retry, degrade, or stop and say so. Never let it continue silently with missing context.

Treat agent state as hot data. Memory and session state are read on every turn. As we described in our guide to agent memory as data storage, that store needs governance; it also needs performance planning.

What to do now

  1. Trace one production agent end to end and measure model time against data path time.
  2. Separate agent reads from transactional traffic with replicas, caches or read stores.
  3. Give agent identities their own rate limits and connection pools.
  4. Co-locate index, data and model endpoint in one region where custody allows.
  5. Define behavior on timeouts so missing context produces an explicit failure, not a weaker answer.
  6. Check regional availability before planning around new instance types such as Z4M, which is in Preview.

The bottom line

Faster models will keep arriving, and they will not fix an agent that waits on a congested database. Treat the data path as part of the agent’s architecture: measure it per step, give agent traffic its own lane, and make timeouts visible instead of letting them quietly lower the quality of answers.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case