← All insights

News analysisLens: United States4 min read

Private AI infrastructure is being sold by workload, not by GPU count. Lenovo AI Express shows the shift

Lenovo and NVIDIA now package hybrid AI capacity in small, medium and large tiers defined by users, tokens per second and model size, shipping in 15 to 25 business days. Agents change the sizing math.

Listen to this article · 6 min

AI-generated narration of the full article.

Long rows of stacked cardboard boxes in a warehouse, in La Madre duotone, beside the words Size the workload
Photo: U.S. Department of Agriculture (rawpixel, CC0)

Buying private AI infrastructure used to look like an HPC project: architects choosing GPUs, interconnects and storage, months of design, and a cluster that was either too big for the first year or too small for the second. Lenovo’s newest offer reads more like buying a validated capacity tier.

What Lenovo announced

Lenovo AI Express, developed with NVIDIA as a quick-start path into Lenovo’s Hybrid AI Factory, comes in three predefined configurations. Lenovo describes each one by the workload it is meant to serve, not by its parts list first:

Tier Lenovo’s workload description Hardware Order to ship
Small Focused inference for tens of users at 30+ tokens per second, models of roughly 7B to 70B parameters ThinkSystem SR650a V4 with two NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs From 15 business days
Medium Higher-throughput inference and agentic AI for hundreds of users at 30+ tokens per second, models of roughly 70B to 400B ThinkSystem SR675 V3 with eight RTX PRO 6000 Blackwell Server Edition GPUs From 20 business days
Large Full-scale generative AI for thousands of users, models up to a trillion parameters ThinkSystem SR680a V4 with NVIDIA HGX B300 From 25 business days

The software options are NVIDIA AI Enterprise or Red Hat AI Factory with NVIDIA, with Veeam Kasten for data resilience. Services cover deployment, ROI validation and GPU optimization, plus a new Premier Support Plus tier for servers. The shipping times apply to eligible configurations after validation, and Lenovo says availability varies by region.

Lenovo also cites its own CIO survey on expected returns from AI. That is vendor research and should not be read as independent evidence.

Why the unit of purchase matters

Describing capacity as “users, tokens per second and model size” is the right move. It is the language of the workload, and it forces the buyer to answer the questions that actually determine whether a system is fast enough: how many people, how quickly they need responses, and how large a model the task needs.

There is a catch, and it is getting bigger: agents break the “users” number.

Sizing from the use case, not from the GPU01Use caseand itslatencytarget02Concurrentusers atpeak03Model callsper task04Tokens percall, inand out05Model sizethe taskneeds06Capacitytier, withgrowthBusiness inputsWhere agents multiply the load
  1. Use case and its latency target
  2. Concurrent users at peak
  3. Model calls per task
  4. Tokens per call, in and out
  5. Model size the task needs
  6. Capacity tier, with growth

Business inputsWhere agents multiply the load

A chat user makes one call per question. An agent may make dozens per task, often with long context.

A person chatting with an assistant generates one model call per question. An agent completing a task can plan, call tools, retrieve, reflect and retry, which can mean dozens of calls per task, each carrying a long context. A tier sized for “hundreds of users” of chat may serve far fewer users of an agentic workflow. Background agents, which run without a person waiting, add load that no “user” count captures.

So the sizing exercise has to start one level lower:

  • Calls per task for each agent, measured in a pilot, not guessed.
  • Tokens per call, input and output, including retrieved context and tool results.
  • Concurrency of agents, including scheduled and event-driven ones.
  • Latency targets per step, because an overnight batch and a customer-facing reply have very different needs.
  • Model mix. If most steps can use a smaller model and only some need a large one, as the routing-by-task argument suggests, the tier you need may be smaller than the largest model implies.

When a packaged tier makes sense

A validated tier with short lead times fits organizations that have already decided some inference must run on infrastructure they control, for the custody and sovereignty reasons we covered in our analysis of the sovereign AI stack, and want to avoid a bespoke design project. It sits one step above the desk-side option NVIDIA just announced with the 64GB DGX Spark: shared, data-center grade, operated by IT.

It does not remove the operating work. Someone still patches the stack, manages models, monitors utilization and plans the upgrade. A packaged tier makes buying faster; it does not make running cheaper.

What to do now

  1. Measure before you buy. Run the target agents in a pilot and record calls per task, tokens per call and latency per step.
  2. Translate agents into load, not into users, and include background agents.
  3. Size for the model mix, not for the largest model you might ever want.
  4. Plan the second year: utilization targets, the growth path between tiers and who owns capacity decisions.
  5. Check regional eligibility and lead times with Lenovo or its partners; the 15-to-25-day figures apply to eligible configurations.

The bottom line

Lenovo AI Express is a sign that private AI infrastructure is becoming a product bought by workload. That is progress. The buyers who get it right will be the ones who know their workload in agent terms, calls, tokens and concurrency, before they pick a tier.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case