← All insights

News analysisLens: United States4 min read

Private AI is getting smaller. When does an enterprise agent belong on a box under the desk?

NVIDIA's 64GB DGX Spark arrives from six OEMs on October 23, sized for local agents with models up to 100 billion parameters. A framework for deciding what should run locally and what should not.

Listen to this article · 6 min

AI-generated narration of the full article.

A small bonsai tree in a shallow pot against a plain gray wall, in La Madre duotone, beside the words Private AI, on the desk
Photo: rawpixel (CC0)

For most enterprises, “private AI” has meant one of two things: a model endpoint in your own cloud tenant, or a rack of GPUs in a data center you control. NVIDIA is pushing a third option further into the mainstream: a capable AI system that sits on a desk.

What NVIDIA announced

On October 2, NVIDIA announced a 64GB configuration of DGX Spark, starting at $4,999, built on the GB10 Grace Blackwell Superchip with a ConnectX-7 network interface. NVIDIA says a single unit runs models of up to 100 billion parameters. The systems will be available from Acer, ASUS, Dell, Gigabyte, HP and MSI starting October 23, so they are announced, not yet shipping.

The second piece is Sync Cluster Assistant. Two 64GB units connected with a QSFP cable pool their memory to 128GB and support models of up to 200 billion parameters; the assistant detects the connected units, validates the configuration and sets up the network without manual work.

The software stack includes NVIDIA’s Agent Toolkit, CUDA-X AI libraries and Nemotron open models, with support for Ollama, vLLM and PyTorch. NVIDIA positions the device for running “capable local agents on device, privately, without cloud dependency,” so developers can work with proprietary data locally before scaling.

NVIDIA’s post does not list availability by country, and the $4,999 figure is a U.S. starting price. OEM pricing and configurations will vary.

The real question: what belongs local?

Where an enterprise agent can runPROVIDER OPERATESYOU OPERATE01Public model API02Model endpoint in yourcloud tenant03On-premises AI cluster04Desk-side system likeDGX SparkElastic, managed, data leaves your hardwareControlled, shared, capital intensive
  1. Public model API
  2. Model endpoint in your cloud tenant
  3. On-premises AI cluster
  4. Desk-side system like DGX Spark
  • Provider operates: Public model API · Model endpoint in your cloud tenant
  • You operate: On-premises AI cluster · Desk-side system like DGX Spark

Elastic, managed, data leaves your hardwareControlled, shared, capital intensive

Each step to the right gains control and loses elasticity and central management.

Local hardware is not a cheaper cloud. It is a different trade. These are the factors that decide it:

Data that should not leave. Source code under export control, unreleased financials, clinical data, M&A material. If the policy answer to “can this go to a hosted model?” is no or takes months, local inference turns a blocked use case into a possible one. This is the custody logic we described in the sovereign stack piece, now at the scale of one team.

Disconnected or constrained environments. Plants, labs, ships and field sites with unreliable connectivity, or air-gapped networks, need inference that does not depend on a round trip.

Development with real data. NVIDIA’s own pitch is development: build and test an agent against proprietary data locally, then decide where it runs in production. That is a sensible use, and it shortens the time security reviews take for a prototype.

Model-size ceilings. 100 billion parameters on one unit, 200 billion on two, covers capable open models, but not the largest frontier models. Some tasks will simply be better on a hosted frontier model. The GPT-6 guide’s logic of routing by task applies here too: local for the steps that must stay private, hosted for the steps that need the most capability, if policy allows.

The costs that do not appear on the price tag

Management. A box under a desk is an unmanaged server unless you make it otherwise. It needs asset tracking, OS and driver patching, disk encryption, endpoint protection, backup of anything that matters, and a decommissioning process for the data on it.

Model supply chain. Local means you choose and update the open models. Somebody has to vet model sources, track versions and re-run evaluations when a model is updated, the same discipline IBM Bob’s self-hosted deployment requires.

Access control. In the cloud, identity and logging come with the endpoint. Locally, you have to decide who can use the system, and how prompts, outputs and agent actions are logged.

Utilization. A desk-side system used by one person a few hours a day is expensive per token compared with a shared endpoint. It earns its keep through privacy, latency or availability, not through unit cost.

What to do now

  1. List the use cases blocked today by data policy, not by technology. Those are the candidates for local inference.
  2. Treat local AI systems as servers in your asset and security processes, from the day they arrive.
  3. Decide model governance before purchase: which open models are allowed, who approves updates, how evaluations are rerun.
  4. Pilot in development first, as NVIDIA suggests, and keep a written decision on where each agent runs in production.
  5. Compare total cost per useful task, including management time, against a private cloud endpoint.

The bottom line

The architecture choice is no longer cloud or data center. A small, capable, private option now sits in between, and it fits a specific set of problems: data that cannot leave, sites that cannot rely on connectivity, and development with real data. Used for those, it can unblock work. Treated as a gadget, it becomes the next unmanaged server holding sensitive data.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case