An HR agent should not need a second copy of Workday. Federation moves the risk to the connection
Databricks' Workday federation, in beta, and Google's borderless lakehouse both read Workday data in place. That removes the copy and makes the integration user and query scope the real controls.
Listen to this article · 5 min
AI-generated narration of the full article.

How many copies of your payroll data exist today? The honest answer in most large companies includes the warehouse extract, the workforce analytics mart, a spreadsheet someone exported for a reorganization, and a backup of each. AI agents threaten to add more: an index here, a cache there, a memory store nobody planned.
Two announcements on October 8 pointed the other way. Databricks released Workday Data Connect federation in beta, which lets Unity Catalog query Workday’s HR and finance tables without ingesting them. The same day, Google said Gemini can read “directly from Salesforce Data 360, SAP, and Workday without copying data”, without giving a release status. When two platforms converge on the same pattern for the same system of record, it is worth understanding what the pattern changes.
How the Databricks connection works
On the Workday side, an administrator shares approved HR and finance tables through Workday Data Cloud and gives a Workday Integration System User, the ISU, read access to them. On the Databricks side, an administrator creates one OAuth connection and a foreign catalog in Unity Catalog. Unity Catalog reads the table metadata from Workday’s Iceberg REST catalog, and Databricks compute reads the data from Workday’s managed storage. Databricks says the connector “does not ingest or copy data into Databricks”. Access is read-only, and Workday remains the system of record.
From there, the tables behave like any other governed data: catalog, schema and table permissions, lineage and auditing in Unity Catalog, and natural-language questions through Genie, such as headcount growth by cost center or attrition by region and job family.
It requires a Unity Catalog workspace with Databricks Runtime 19 or later, Workday Data Connect enabled on the tenant, and a workspace administrator turning on the beta.
| Question | Copy into the platform | Federate in place |
|---|---|---|
| Where the data rests | A second store you must secure | Workday's managed storage |
| Who decides what is visible | The pipeline and warehouse grants | Workday sharing and the ISU role, then Unity Catalog grants |
| Deletion and retention | Must be repeated in the copy | Follow the source |
| Freshness | As of the last load | Current, with no stated guarantee |
| Main cost | Storage and pipelines | Compute per query |
The integration user becomes the ceiling
Federation concentrates authority in one place: the integration user’s read role. Whatever the ISU can read is the most that any analyst, dashboard or agent downstream can ever see through this connection. That makes the ISU’s role the most important access decision in the design, and it should be reviewed as such, with HR and privacy in the room.
Below the ceiling, the post describes permissions in Unity Catalog, at catalog, schema and table level. It does not say how Workday’s own per-person security, such as a manager seeing only their reports, carries over. Until Databricks documents that, plan as if Unity Catalog grants, row filters and column masks are the only per-user controls downstream. Compensation, performance ratings and medical leave need masks before anyone points a natural-language agent at the catalog, because a question like “who are the highest-paid people in this team?” is easy to ask.
No copy is not no copies
Federation removes the bulk copy. It does not remove the smaller ones that agents create: query results, Genie answers shared in a channel, an agent’s memory of last week’s attrition analysis, logs of prompts that contain names. Those are exactly the copies we described when we argued that custody is the first question. They need retention rules of their own.
Two limits are worth writing down before the first pilot. Data still moves for processing: Databricks compute reads from Workday’s storage, so where the workspace runs matters for where personal data is processed. And federation reads the current state. Questions about trends over years need history, which the post does not address and which may still require snapshots.
Federation is the right default for systems of record, and Workday is a good place to start because the cost of an extra copy of HR data is so high. It changes the review rather than ending it: from “how do we secure the copy?” to “how narrow is the connection, and who checks it?”