← All insights

News analysisLens: United States3 min read

Is the agent the product? Copilot Studio's October update suggests the product is the whole process

Microsoft now builds apps, deterministic workflows and agents in one Copilot Studio environment, and extends evaluation from agent conversations to automation. Approval has to follow.

Listen to this article · 5 min

AI-generated narration of the full article.

A drafting compass resting on technical blueprints, in La Madre duotone, beside the words Test the whole process
Photo: MichaelGaida (StockSnap, CC0)

When an enterprise approves an AI agent, what exactly is it approving?

The usual answer treats the agent as the unit. It gets a design review, an evaluation set, a security sign-off and a release. Microsoft’s latest Copilot Studio update, published on October 7 by Jason Moore, vice president of product, puts pressure on that assumption, because the thing it helps teams build is no longer an agent. It is a process with an agent inside it.

Three kinds of logic in one place

The update’s central idea is that makers can “build apps, workflows, and agents in one place”. Microsoft’s own example is employee onboarding: an app gives managers and employees a structured interface, workflows execute the repeatable onboarding steps, and an agent answers questions. In the post’s words, teams can bring together “the interface people use, deterministic workflows the business may rely on, and agents that reason, adapt, and act”.

The pieces ship at different maturities:

  • Apps, generated from a description and refined in conversation with access to the code, are in public preview.
  • Copilot Managed Runtime, which we analyzed when it was announced as the hard part of operating AI-built apps, is still in public preview. It now covers apps from Copilot Studio, Copilot Cowork and Copilot Code, with Microsoft-operated hosting, identity, governed data access, lifecycle management and Git-backed versioning.
  • Hooks for agents, through the GitHub Copilot harness, are in preview. They run deterministic logic at points in an agent’s lifecycle such as session start, tool calls and errors. Microsoft’s phrasing is precise: “While a workflow defines the logic to execute, a hook determines when that logic runs.”
  • Foundry IQ integration, for knowledge bases reused across agents with private networking, is generally available.
  • The Review panel, which surfaces blockers and warnings, including evaluation warnings, before publishing, is generally available.

Where the agent-as-product view breaks

If the solution is an app, a set of workflows and an agent, many failures will not live inside any one of them. They will live at the joins. An agent passes the wrong employee ID to a workflow. A workflow returns an error and the agent tells the user it succeeded. An app shows the agent’s draft as if it were the approved record.

An evaluation that only looks at the agent’s conversation misses all three. That is why the most consequential line in the update is a short one: evaluations in Copilot Studio “are expanding beyond agent conversations to automation”.

The new graders map onto those joins. Task Completion checks whether the intended outcome was reached, not whether the reply sounded right. Tool Accuracy checks whether the agent chose the expected tools and passed the expected inputs, which is exactly where an agent hands off to a deterministic workflow. Safety and latency checks sit alongside them, and a library of custom graders can be reused across agents. Microsoft does not state a release status for each of these evaluation features, so check them in your tenant before planning around them.

Deterministic where the business relies on it

The update also confirms a design rule we drew from Snowflake’s accounts payable system: keep the deterministic parts deterministic. Hooks are a way to enforce it from inside an agent’s run. A check at every tool call can validate parameters against policy before a workflow executes. A hook on errors can send a failure to a defined path instead of letting the agent improvise an explanation. Both are design choices teams will make, not defaults Microsoft provides.

One more detail deserves attention: a new Evaluation Viewer role lets reviewers and other stakeholders see evaluation results. That puts the evidence in front of the person who owns the business process, not only the person who built the agent. As we argued about agent observability beyond uptime, a system that returns a success code can still be wrong, and the people best placed to spot it are the ones who know what right looks like.

The practical change is in what gets approved. Test cases should start where a person starts, in the app, and end where the business record changes, in the system the workflow writes to. Tool accuracy should be graded on every workflow the agent can call. And production commitments should rest on the parts that are generally available today, not on the previews around them. The agent is a component. What the business approves, and what an auditor will eventually sample, is the process.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case