Your company no longer has an AI model. It has a portfolio, and it needs managing like one
Specialist models, decision models, domain operators and routing defaults arrived in one week. The single-model era is ending inside the enterprise. A framework for the six roles in a model portfolio.
Listen to this article · 7 min
AI-generated narration of the full article.

For two years, the enterprise question was “which model should we use?” Look at what shipped in the past week and the question no longer makes sense.
OpenAI started billing for GPT-Rosalind, a life sciences model available only for approved research. Cloudflare released decision models that return probabilities instead of text. OpenAI and Synopsys announced a model trained to operate chip design tools. Chatham Financial described a platform where four OpenAI models split the work by difficulty, and Albertsons described predictive models and generative AI working together. Add the routing guidance OpenAI published with GPT-6, and the pattern is hard to miss.
Our view, and it is a view rather than a fact: the single-model era is ending inside the enterprise. What replaces it is a portfolio of models with different jobs, owners, costs and rules. Companies that keep managing “the model” as one thing will either overpay, under-control, or both.
The six roles in a model portfolio
Not every company needs all six. Most production AI estates will have at least four within a year.
- Frontier generalist. Planning, open-ended reasoning, hard synthesis. Expensive per call, worth it on the steps that need it. In our analysis of OpenAI’s GPT-6 guide, this is the model you route to on purpose, not by default.
- Efficient default. The cheaper general model most traffic should hit first. Chatham’s internal builder platform makes GPT-5.6 Terra the default and lets each app opt into Sol. That default is a governance decision written in configuration.
- Domain specialist. Models trained for a field, like GPT-Rosalind for biology and drug discovery. They often come with restricted access tied to an approved use, which makes them governed assets, as we described in today’s analysis of GPT-Rosalind’s billing.
- Domain operator. Models trained to drive a specialist tool, like GPT-Synopsys for electronic design automation. Still early, but the direction is clear, and it raises the custody stakes because the tool’s data travels with the model.
- Decision model. Small, fast models that answer typed questions with probabilities: route, block, approve, escalate. Cloudflare’s Clef is the clearest example. They take the frequent, narrow questions off the expensive models.
- The predictive models you already own. Forecasting, propensity and risk models did not become obsolete. Albertsons describes combining predictive models with generative AI to produce explainable recommendations. In many companies, the most valuable generative feature will sit on top of an existing model, not replace it.
Retrieval and embedding models belong in the inventory too, and so do the deterministic verifiers that are not models at all: test suites, reconciliations, validators. They are the authority the portfolio answers to, as GPT-Synopsys’s design makes explicit.
- Decision model classifies the request
- Efficient default handles routine work
- Frontier or specialist on explicit route
- Existing predictive model supplies scores
- Deterministic check verifies
- Person decides exceptions
- Route: Decision model classifies the request
- Work: Efficient default handles routine work · Frontier or specialist on explicit route · Existing predictive model supplies scores
- Authority: Deterministic check verifies · Person decides exceptions
Fast, typed decisionModels with different jobs and costs
What changes when you manage a portfolio
Routing becomes policy. Which tasks go to which model, with what default and what upgrade path, belongs in one place (the gateway or orchestration layer), versioned and logged. Chatham’s default-plus-upgrade rule is a good template.
Every model needs an owner and an evaluation set. A portfolio of six models with one shared benchmark is not governed. Each role needs its own measure of success: calibration for decision models, domain accuracy for specialists, cost per successful task for the default.
Lifecycle multiplies. Each model in the portfolio will be retired on its own schedule. We argued that model retirement is the new API deprecation; with six models, that becomes a calendar, not an event.
Access differs by model. Some models are open to everyone; some, like trusted-access specialists, are restricted to approved uses and named teams. The access model has to be per model, not per platform.
Cost is measured per outcome. The portfolio exists to put each task on the cheapest model that does it well enough. That only works if you measure accepted outputs, not tokens.
For U.S. banks and their suppliers, there is a familiar frame: the Federal Reserve’s SR 11-7 already expects a firm-wide model inventory with owners, validation and ongoing monitoring. A generative model portfolio is new technology inside an old discipline. The firms that extend their existing inventory instead of starting a parallel “AI register” will move faster at review time.
In a Microsoft estate, Foundry’s model catalog and model router make the portfolio concrete. The tools help; the routing policy, ownership and evaluation still have to be decided by the company.
What would prove this view wrong
If one frontier model became cheap, fast and accurate enough across specialist domains that routing stopped paying off, the portfolio would collapse back into a single model. We do not expect that in the next year, because specialists are arriving with access rules and tool integrations a generalist cannot replicate by being smarter. But it is the signal to watch.
A one-page portfolio review
- List every model in production or pilot, including embeddings, decision models and existing predictive models.
- Assign each a role from the six above, an owner and an approved use.
- Write the routing policy: default, explicit upgrades, restricted specialists.
- Define one success measure per role and one evaluation set per model.
- Put retirement dates on a calendar and attach a regression set to each.
- Report cost per accepted output by use case, not total token spend.
The bottom line
The useful question is no longer which model is best. It is which job each model does, who owns it, how it is evaluated and what decides whether its output is right. Companies that answer those four questions per model will get the benefit of specialists without losing control of the whole.