Who holds what? Custody is becoming the first question in enterprise AI architecture
A month of announcements points the same way: agents create new copies of sensitive data in sessions, snapshots, traces and indexes. Why a custody map should come before the security review.
Listen to this article · 5 min
AI-generated narration of the full article.

In the first year of enterprise generative AI, the opening question was “which model?” In the second, it became “which platform?” Looking back at what we covered over the last month, a different question now comes first in almost every serious announcement: who holds what?
This is our view of why that happened, and what enterprise teams should do about it.
The pattern behind a month of announcements
Read side by side, the announcements we analyzed in September and early October describe the same shift from different angles:
- IBM made its agentic development platform available on premises and air-gapped, so the agent and the models can sit inside the customer’s boundary. We covered it in our analysis of self-hosted IBM Bob.
- The same Claude model now arrives through paths where either the model provider or the customer’s cloud provider runs inference, and some tools, such as MCP connections, sit outside zero data retention. See our piece on Claude’s delivery paths.
- Anthropic’s Enterprise Frontier Safeguards split safety monitoring: the provider detects misuse, the customer holds the data, keys and review. We called it a custody split.
- Google’s Agent Substrate keeps idle agents as snapshots of their memory and files in storage, as we described in our look at agent fleets.
- Snowflake’s agent observability stores prompts, completions, retrieved documents and tool calls as traces, which we argued in our article on observability makes telemetry a sensitive data set.
None of these is mainly about model quality. Each one is about where data and control live, and who operates them.
Why custody moved to the front
The reason is structural. A chat assistant that answers a question creates little new data. An agent that plans, retrieves, calls tools, pauses and resumes creates new copies of sensitive information in places that did not exist before: session stores, memory snapshots, vector indexes, trace stores, caches, tool exchanges. Most enterprise data inventories list none of them.
Security and privacy teams have noticed. In regulated environments, the first questions in an AI review are rarely about the model. They are about where the data goes, who can read it later and how long it stays.
- Inference
- Prompts and outputs
- Agent state: sessions, snapshots
- Retrieval index and embeddings
- Traces and evals
- Tool calls leaving the boundary
New data stores most inventories do not list yet
The custody map: a simple discipline
Our recommendation is to make a custody map part of every AI system design, before the security review rather than in response to it. For each of the six points above, answer five questions:
- Where does it live? Which service, which account, which region.
- Who operates it? The model provider, your cloud provider, a SaaS vendor or your own team.
- Who can read it? Including support staff, administrators and other agents.
- How long is it kept? And how it is deleted.
- Under which jurisdiction? Which matters as soon as personal or regulated data is involved.
Three principles make the map useful:
Derived data is still data. Embeddings, traces and snapshots are derived from documents, conversations and records. Treat them with the classification of their source, not as technical byproducts.
Choose topology per data class, not per vendor. The same organization may run an internal assistant through a SaaS path and a claims agent fully inside its own cloud. The custody map is what justifies the difference.
Custody includes evidence. Audit logs, approvals and evaluation results are records someone will ask for later. Decide where they live with the same care as the data itself.
What this means in regulated industries
In life sciences, the logic will be familiar: data integrity expectations such as ALCOA+ already require records to be attributable, original and retained. If an agent participates in a regulated process, its traces and approvals may become part of that record. In US banking, third-party risk management guidance from the federal banking agencies already asks institutions to understand where their data sits with each provider. AI agents add more providers and more places to that list, not a new kind of question.
What we expect next
Our expectation, and it is a view rather than a fact: within the next year, enterprise buyers will ask AI vendors for a custody map the way they now ask for security reports, and vendors will start publishing them. The announcements of the last month already read like early drafts.
Teams that build the map now will move through procurement and security review faster, and will have an easier time choosing between SaaS, cloud-hosted and self-hosted options as those options keep multiplying. The rest of the controls are in our 2026 enterprise AI stack guide.