← All insights

News analysisLens: United States3 min read

An AI agent can return a 200 and still be wrong. Snowflake's Agent Observability measures what uptime can't

Snowflake is bringing agent tracing, cost signals and LLM-as-judge evaluation to Observe. Why production AI monitoring must cover behavior, quality and cost, and why the telemetry is sensitive too.

Listen to this article · 5 min

AI-generated narration of the full article.

A dense aircraft cockpit panel full of analog gauges, in La Madre duotone, beside the words Beyond uptime
Photo: rawpixel (CC0)

Traditional monitoring asks whether a service is up, fast and free of errors. An AI agent can pass all three checks and still fail: it answers with confidence from the wrong document, calls the right tool with the wrong parameters, or loops through twenty model calls to finish a task that should take three. Nothing in a standard dashboard turns red.

Snowflake’s announcement of Agent Observability, coming to its Observe product, is built around that gap. It is a useful reference for what production AI monitoring now needs to cover, whatever platform you use.

What Snowflake announced

According to Snowflake, Agent Observability is coming soon to private preview, with access requested through Observe account teams and a launch event planned for October 22. The capabilities it describes:

  • Traces of agent behavior: prompts, completions, retrievals, tool calls, token usage and session identifiers, searchable across traces, sessions and conversations.
  • Performance and cost signals derived from those traces: latency, errors, token usage and estimated cost, with comparisons across agents, models and workflows.
  • Quality evaluation: online LLM-as-judge evaluations on production traffic to surface hallucinations and quality issues, plus tracking of offline evaluation results.
  • Open standards: an SDK that follows the emerging OpenTelemetry GenAI semantic conventions, with Apache Iceberg for data access. Supported frameworks include LangChain, the OpenAI Agents SDK and Anthropic’s agent SDK, with custom frameworks through an OTLP endpoint.

As with any preview, the scope, pricing and limits can change before general availability.

Four questions production AI monitoring has to answer01Is it up? Uptime,errors, latency02What did it do?Prompts, retrievals,tool calls03Was it right? Onlineand offline evals04Was it worth it?Tokens and cost pertaskSignals specific to AI agents
  1. Is it up? Uptime, errors, latency
  2. What did it do? Prompts, retrievals, tool calls
  3. Was it right? Online and offline evals
  4. Was it worth it? Tokens and cost per task

Signals specific to AI agents

Classic monitoring stops at the first box. An agent can pass it and fail the other three.

What changes for enterprise teams

Quality becomes a runtime signal. Most teams evaluate an agent before launch and then watch only operational metrics. Online evaluation on live traffic moves quality into the same loop as latency. That is the right direction, with one caveat: an LLM judge is itself a model, and it needs its own calibration against human review before its scores drive decisions.

Cost becomes visible per behavior, not per invoice. Token usage tied to a trace shows which agent, workflow or model choice is expensive, and why. That is the level at which cost can actually be managed.

Open conventions reduce lock-in. Instrumenting agents with OpenTelemetry GenAI conventions means the same traces can feed a different backend later. Ask that of any observability vendor.

Telemetry is sensitive data. This is the part most announcements skip. A trace that stores prompts, completions and retrieved documents holds exactly the information the agent was trusted with: customer details, contract terms, health or financial data. The observability store needs the same access controls, masking, retention rules and residency decisions as the source systems. Otherwise it becomes the easiest place to read everything the agent ever saw.

We argued in our analysis of Palantir’s AIP Evolve that evaluation is becoming the control surface for AI systems. Observability is where those evaluations meet real traffic.

For US enterprises

  1. Define what “wrong” means per agent before buying tooling: wrong answer, wrong tool call, policy breach, excessive cost. Each needs its own signal.
  2. Instrument with open conventions so traces are portable across tools.
  3. Classify the trace store like the most sensitive system the agent reads from, and route access through your existing identity controls.
  4. Calibrate LLM judges on a sample of human-reviewed cases, and recheck when models change.
  5. Set cost budgets per agent and per task, and alert on drift, not just on totals.

The bottom line

Agent Observability is one more sign that operating AI in production is becoming its own discipline. Uptime is necessary and no longer sufficient. The teams that can answer what an agent did, whether it was right and what it cost, per task, will be the ones allowed to scale it. The full list of controls is in our 2026 enterprise AI stack guide.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case