AI spend now needs guardrails, just like permissions. Platforms are shipping them, and CFOs want a say
Snowflake made per-user AI quotas generally available, Google added spend caps for agents, and IBM reports CFOs taking a larger AI role. How to design cost controls before agents scale.
Listen to this article · 6 min
AI-generated narration of the full article.

A chat assistant has a fairly predictable cost: people type, the model answers. An agent does not. It plans, retries, calls tools, reads documents and loops until it decides it is done. Multiply that by every employee with access, or by thousands of automated runs a day, and AI spend stops being a line item and becomes a variable that can move overnight.
Three announcements from the last few weeks show the market responding from two directions: platforms are turning cost control into product features, and finance leaders are taking a larger role in AI decisions.
Platforms ship spend controls
Snowflake made per-user quotas for AI generally available in September. Administrators can set daily and monthly limits per user, separately for each governed service: Cortex AI Functions, Cortex Agents, its developer assistant and warehouse compute. When a user hits a limit, Snowflake says access to that specific service pauses until the next cycle, without affecting other services or other users. Limits can be set in SQL, in the Snowsight interface or in natural language.
Google Cloud added a set of FinOps controls for agents in August. Among them: project-level spend caps that temporarily pause an agent’s API calls when the monthly limit is reached, with alerts at 50%, 80% and 100%; anomaly detection that flags unusual AI spend and names the top three SKUs driving it; pooled quotas with admin-controlled overages for its Antigravity developer platform; and a deferred execution option, announced as coming soon, that runs eligible agent work off-peak at up to half the inference cost.
The pattern: spend limits are being designed like permissions, scoped to a user, a project or a service, and enforced by the platform.
Finance moves closer to AI decisions
On September 30, IBM’s Institute for Business Value published a study of 1,500 CFOs across 33 geographies and 26 industries, surveyed between February and April 2026. According to IBM, 62% report an expanded role in leading enterprise technology or AI strategy, and 56% report greater authority over portfolio management and capital reallocation. Only 6% say their finance function has reached a transformation-ready state.
These are IBM’s survey findings, not independent measurements. But the direction matches what the platform announcements imply: when AI cost becomes variable and material, the people who own budgets want visibility and control, not a monthly surprise.
- Per-user quotas
- Per-project caps
- Anomaly alerts
- Spend attributed per agent and task
- Off-peak execution for work that can wait
- Commitment discounts
- Limit: Per-user quotas · Per-project caps
- Detect: Anomaly alerts · Spend attributed per agent and task
- Optimize: Off-peak execution for work that can wait · Commitment discounts
Can pause work when reached: design per workloadLower unit cost, no change to behavior
The design question nobody asks: what happens at the limit?
A hard limit is a safety feature for an employee experimenting with an assistant. It is an outage for a customer-facing agent. Both Snowflake and Google describe pausing work when a limit is reached. That is the right default for exploration and the wrong one for a production workflow that customers depend on.
So cost guardrails need the same design discipline as security controls:
Tier workloads by criticality. Exploration and individual productivity get hard per-user limits. Production agents get budgets with alerts, controlled overage and a named owner who decides what happens next.
Attribute spend to behavior. A project total tells you that you overspent. Cost per agent, per task and per model tells you why. That requires the kind of tracing we discussed in our analysis of agent observability.
Separate work that can wait. Batch enrichment, nightly analysis and report generation can often run off-peak at lower cost. Interactive work cannot. Deciding which is which is an architecture choice, not a billing setting.
Give finance a real interface. If CFOs are taking on AI portfolio decisions, they need spend data organized by use case and outcome, not by SKU. Building that view is easier at design time than after a surprise invoice.
For US enterprises
- Inventory where AI spend can originate: assistants, agents, AI functions inside data platforms, developer tools.
- Set per-user limits for exploration on every platform that supports them.
- Give each production agent a budget, an alert threshold and an owner, and agree in advance what happens when it is exceeded.
- Turn on anomaly detection and route it to the owner, not just to a shared FinOps inbox.
- Report spend by use case to finance, alongside the value each use case is expected to deliver.
The bottom line
The cost of AI agents is variable by design. Platforms now provide the levers: quotas, caps, alerts and cheaper ways to run work that can wait. Using them well is an operating-model decision about which workloads can be paused, who owns each budget and how finance sees value. In our 2026 enterprise AI stack guide, that decision belongs to the layer nobody ships: operating ownership. Cost is often what stops a program that is otherwise working.