← All insights

News analysisLens: United States4 min read

Local or cloud is no longer the question. Windows and NVIDIA turn model placement into a routing policy

Microsoft is building Windows around hybrid intelligence and NVIDIA is putting large local models on desks. The decision that matters is who writes the rules for routing each task.

Listen to this article · 6 min

AI-generated narration of the full article.

A dirt road splitting into two paths between green fields, in La Madre duotone, beside the words Route each task
Photo: rawpixel (CC0)

For two years the enterprise AI question on the desktop was binary: send the work to a cloud model, or do without. On October 7, Microsoft and NVIDIA made a coordinated case that the answer is now “both, per task”, and that the operating system should help decide.

That shifts the real decision. It is no longer where your AI runs. It is who writes the rule that sends each task somewhere.

What was announced

Microsoft described Windows as a platform for what it calls hybrid intelligence: agents run locally when that makes sense and reach the cloud when they need to. The pieces, with their status:

  • GitHub HydraFusion routes each task to the right model, in the cloud or on the device. Microsoft lists it as an experimental preview later in October.
  • Windows ML runs local models on the GPU, NPU or CPU.
  • Microsoft Execution Containers, generally available, set which files and networks local agents can reach. We covered them in our analysis of agent containment on Windows.
  • Hybrid features in Copilot are expected to roll out in the coming months, opt-in, starting in select markets.

Microsoft also says 40% of business laptops are now Copilot+ PCs and that those PCs run about 2 trillion local inferences a month. Those are Microsoft’s own figures.

NVIDIA introduced RTX Spark, which pairs a Blackwell RTX GPU with up to 6,144 cores, a Grace CPU with up to 20 cores and up to 128GB of unified memory, rated at one petaflop of FP4 performance. Laptops from Acer, ASUS, Dell, HP, Lenovo, Microsoft, MSI and Gigabyte opened for preorder on October 7, with shipping from October 16; compact desktops follow in November. Microsoft’s Surface RTX Spark Dev Box ships in November, in the U.S. only. Neither company listed availability by country for the rest.

NVIDIA also showed DGX Station for Windows as a preview: 748GB of coherent memory and up to 20 petaflops of FP4 compute, which NVIDIA says is enough for models up to trillion-parameter scale. Microsoft says it will be available later this year.

What actually changes

Local inference is not new; we asked when an enterprise agent belongs on a box under the desk a few days ago. What is new is that routing is becoming a platform feature. When the operating system and the developer tools decide which model handles which step, placement stops being an architecture review held once and becomes a runtime decision made thousands of times a day.

That decision has at least four inputs, and each one belongs to a different owner:

  • Privacy: can this data leave the device at all? Security and privacy own that answer.
  • Latency: does the user wait for this step? Product teams own that.
  • Cost: what does a cloud call cost against hardware already paid for? Finance and FinOps own that.
  • Capability: is a local model good enough for this task? Only evaluation can tell.
Routing one agent taskROUTING POLICY OWNED BY THE ENTERPRISE01Task arrives02Data policy:may it leavethe device?03Capabilitycheck fromevals04Cost andlatency budget05Local model orcloud modelAgent requestRules with named owners
  1. Task arrives
  2. Data policy: may it leave the device?
  3. Capability check from evals
  4. Cost and latency budget
  5. Local model or cloud model
  • Routing policy owned by the enterprise: Data policy: may it leave the device? · Capability check from evals · Cost and latency budget

Agent requestRules with named owners

A vendor router can execute the decision. The rules it follows should be yours, and every placement should leave a record.

The router is the new control point

A router that picks the “right model” optimizes for something. In a vendor tool that is usually quality and token spend. Your enterprise may need it to optimize for data classification first, which a generic router cannot know.

Data classification has to reach the router. If a file is labeled confidential, the routing decision should see that label before any cloud call. Ask any vendor how its router consumes your labels, not just whether it supports local models.

Local capacity becomes a FinOps line. Paying once for a large-memory workstation that runs a capable open model changes the cost curve for heavy users, like developers running agents all day. It also adds hardware refresh, model updates and patching to someone’s job. We argued in our piece on managing a model portfolio that each model you run is something you operate; a model on every desk is many of them.

Evaluation has to cover both paths. If the same task can land on a 125B local model or a frontier cloud model, quality tests must run on both, or the router will quietly lower quality to save cost.

Previews are previews. HydraFusion is experimental, DGX Station for Windows is a preview, and the Copilot features are months away. Pilot, do not standardize.

What to do now

  1. Write the routing policy before buying the hardware: which data classes must stay local, which tasks need frontier quality, what each path may cost.
  2. Pick one heavy-use group, usually developers, for a local-plus-cloud pilot, and measure cost per task on both paths.
  3. Require a log of each placement decision: task, model, location and reason.
  4. Extend evaluations to the local models you allow, with the same test sets as the cloud ones.
  5. Pair any local agent with containment and decide who manages those devices.

The bottom line

The hybrid runtime is arriving on the desktop, partly shipped and partly in preview. The hardware will sort itself out. The part to get right now is the routing policy: written by the enterprise, enforced by the platform, and visible in a log.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case