← All insights

Trend analysisLens: United States3 min read

Agents now read your business definitions where they live. Next question: whose definition of net margin wins?

Google's Knowledge Catalog reads definitions from Databricks, dbt, LookML and SAP; Databricks made business domains GA. Shared semantics are becoming agent infrastructure, and need owners.

Listen to this article · 5 min

AI-generated narration of the full article.

A coiled tape measure, in La Madre duotone, beside the words Shared meaning
Photo: rawpixel (CC0)

Ask two agents for last quarter’s net margin. One reads the finance team’s LookML model; the other reads a dbt metric an analytics team wrote two years ago for a different purpose. Both answer confidently, a point and a half apart. Neither is hallucinating. They are reading two definitions of the same word.

That problem moved closer to the center of agent architecture on October 8. At Gemini at Work, Google introduced a Knowledge Catalog to “map business definitions once so all agents use them”, for terms such as “net margin” and “addressable market”. Gemini, Google says, reads those definitions “directly where they sit” in Databricks, dbt, LookML or SAP, rather than asking customers to move them. The same day, Databricks made its Discover page and domains generally available: domains and subdomains organize tables, dashboards, notebooks and Genie Agents by business area, so people can find assets without knowing the catalog hierarchy.

Google gave no release status for the Knowledge Catalog. Even so, the direction across both vendors is clear: business meaning is becoming something agents query, not something analysts remember.

Where business meaning sits for agents

Agents and assistants

  • Gemini
  • Genie Agents
  • Other copilots

Shared business definitions

  • Net margin
  • Active customer
  • Headcount

Each term needs one owner, a version and a test

Where definitions are written today

  • dbt
  • LookML
  • Databricks
  • SAP

Data

  • Tables
  • Documents
  • Objects
Reading definitions in place avoids another copy. It does not decide which source is authoritative when two of them disagree.

Reading in place exposes conflicts; it does not resolve them

Reading definitions where they already live is the right design. It is the layer we have argued enterprises should own once across vendors, and it avoids yet another semantic model that drifts from the originals. In our analysis of Genie One for regulated finance, the governed ontology mattered more than the chat interface for the same reason.

But a catalog that reads from four sources inherits the disagreements between them. When dbt and LookML define net margin differently, a federated catalog surfaces two answers or silently prefers one. Neither is acceptable for a board report. The decision about which source is authoritative for each term is a governance decision, and the vendor cannot make it for you.

Definitions need versions, owners and tests

Once agents depend on a definition, changing it changes answers. When finance redefines active customer, every agent that reports on customers shifts at once, including answers already sent to executives last month. That calls for practices data teams apply to code more often than to metrics:

  • One named owner per term, usually in the business, not in the data team.
  • A version and an effective date, so a past answer can be explained with the definition in force when it was given.
  • An evaluation set per critical term, rerun when the definition or the model changes. An answer that is right under one version and wrong under the next is not a model failure; it is a release.

Enrichment is new data, with new questions

Google’s Smart Storage goes a step further. It enriches unstructured objects in place, “writing context back directly onto the object itself”, and Google says it inherits the existing security posture. Google puts unstructured data at ninety percent of enterprise data.

Enrichment written by a model becomes part of the data, and other systems will trust it. Two checks are worth making before turning it on: who reviews the quality of generated context on sensitive documents, and whether metadata written onto an object is visible to everyone who can list the object or only to those who can read it. Those are often different groups.

Databricks’ domains raise a smaller version of the same point. Grouping Genie Agents by business area makes the right agent easier to find, which reduces duplicate agents built because nobody could find the first one. The release note describes domains as organization; access is still granted in Unity Catalog. Findability and permission are different controls, and both are needed.

The semantic layer is becoming the place where a company’s agents agree with each other. That only works if someone owns what each word means.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case