← All insights

News analysisLens: United States4 min read

Stop picking one best model. OpenAI's GPT-6 guide makes the case for routing by task

OpenAI's new guide to the GPT-6 family treats model, reasoning effort and speed as choices per task. For enterprise teams, the unit to optimize becomes cost per successful task.

Listen to this article · 6 min

AI-generated narration of the full article.

An aerial view of a multilevel highway interchange, in La Madre duotone, beside the words Route by task
Photo: rawpixel (CC0)

Many enterprise AI programs still run on a single decision made early: “we use model X.” It made sense when one model was clearly ahead and the price differences were modest. It makes less sense now that a single vendor ships a family of models with different strengths, adjustable reasoning and paid speed tiers, and a single agentic workflow may call a model a dozen times for very different jobs.

OpenAI’s guide to its GPT-6 family, published October 2, reads like a product manual. The useful part for enterprise teams is the production architecture underneath it.

What the guide says

OpenAI frames model choice and reasoning level as “an intelligence/price tradeoff” and maps three models to three kinds of work:

  • GPT-6 Astra for the hardest reasoning, where maximum intelligence is needed.
  • GPT-6.1 Sol for complex coding, research and computer use. It also supports multi-agent workflows in the Responses API, currently in beta.
  • GPT-6 Luna for focused, repeated work at scale with a clear goal, such as extracting invoice fields, classifying requests or producing structured summaries.

On top of the model, teams choose a reasoning effort: low for routine extraction or small edits, medium for work that needs judgment, high for difficult debugging or careful review, and extra high or max only when high falls short and the gain justifies the time and cost. In the API, effort can change mid-conversation without breaking the cache.

Then speed: Fast mode gives faster, more consistent responses at a higher per-token price, and Ultrafast, available for GPT-6 Astra, accelerates token generation at a premium.

The production section is the most practical. Cut context the task does not need. Run independent tasks in parallel. Put stable instructions and reference material first so prompt caching works; OpenAI says cached input tokens cost up to 95% less than uncached ones, depending on the model, and reminds teams to include cache writes and long-context rates in cost estimates. Use compaction for long conversations. And before deploying, run representative tasks and measure task success, latency and cost per successful task.

The architecture lesson: route by step, not by vendor

Put those pieces together and the unit of design is no longer “the model.” It is the step:

One workflow, several routes01Classify therequest02Extract fields03Plan orcompareoptions04Review therisky case05Draft thereplySmall model, low effortMid model, medium effort
  1. Classify the request
  2. Extract fields
  3. Plan or compare options
  4. Review the risky case
  5. Draft the reply

Small model, low effortMid model, medium effort

Each step gets the cheapest route that passes its own evaluation. The hard step pays for intelligence; the rest do not.

That has four consequences.

Routing becomes a component. Someone has to decide, per step, which model, which effort and which speed tier, and change that decision as prices and models move. That logic belongs in a gateway or orchestration layer, not scattered across prompts. We described the gateway as a policy enforcement point in our analysis of C1’s LLM Gateway; routing by task is one of the policies it should carry.

Evaluation decides the route, not reputation. “Use the best model” is a guess. “Use the cheapest route that passes this step’s evaluation set” is a decision you can defend to finance and to an auditor. It also gives you the regression suite you need when a model is retired, which is the problem we covered in our piece on model retirement.

Latency is a budget, not an afterthought. A customer-facing chat step may have a two-second target; an overnight reconciliation has hours. Paying for Fast or Ultrafast makes sense only where a latency objective exists.

Context is a cost line. Caching and compaction are not optimizations for later. Prompt structure (stable first, variable last) determines whether caching works at all, and caching determines whether a high-volume workflow is affordable.

What did not change

The guide is explicit about something easy to skip: decide which actions the model may take independently and which require approval, and replace blanket “always ask” rules with clear boundaries. Routing optimizes cost and quality; it does not settle authority. A cheaper model on a step that can send an email or update a record still needs the same controls as an expensive one.

It is also a single-vendor guide. The same logic applies across vendors, and a routing layer that is vendor-neutral keeps that option open. For organizations standardized on Microsoft, the same pattern applies to model deployments in Microsoft Foundry: per-step routing, evaluation sets and cost per task, whatever models sit behind them.

What to do now

  1. Break each production workflow into steps and write down, per step, the quality bar, latency target and acceptable cost.
  2. Build a small evaluation set per step, not per application.
  3. Measure cost per successful task, including retries, cache writes and human correction time, not cost per token.
  4. Move model names into routing configuration so a change is a config release with test evidence, not a code change.
  5. Review prompt structure for caching: stable instructions and reference material first, task details last.
  6. Keep authority separate from routing: approval rules attach to the action, not to the model.

The bottom line

The question “which model should we standardize on?” is giving way to “which route does each step deserve?” That is a better question, because it can be answered with evidence and revisited every time prices or models change. The teams that answer it per step, with evaluations and a cost-per-success metric, will spend less and break less.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case