Not every agent decision needs a generative model. OpenAI's Decisions API prices the decision layer apart
OpenAI opened its Decisions API in public beta: GPT-6 Luna returns probabilities, choices or scores instead of prose, billed on input only. The decision layer is becoming a product category.
Listen to this article · 4 min
AI-generated narration of the full article.

Earlier this week, Cloudflare released open decision models and we argued that an agent should not ask a frontier model every small question. On October 6, OpenAI made the same architectural argument with its own product, and gave it a price list.
Put the two side by side and the decision layer stops looking like a pattern and starts looking like a product category.
Same idea, different terms
OpenAI’s Decisions API is available to all developers in public beta. It runs on GPT-6 Luna, accepts text and image inputs, and returns one of three structured outputs instead of free text: predicates, a probability that a statement is true; choices, a selection from options you define, with confidence scores; and scores, a number on a range you set.
Cloudflare’s Clef models answer the same kinds of questions. The difference is in custody and billing. Clef’s weights are open and can run anywhere; OpenAI’s model is a hosted API. OpenAI says decisions are up to 10 times faster than calling GPT-6 Luna through the Responses API, and prices them at $0.10 per million input tokens, with no charge for output tokens or for cache reads and writes. Both the speed and the price are OpenAI’s figures for a beta, and the announcement does not mention availability through cloud partners.
The price list changes the design
You pay for context, not answers. With output free and input billed, the cost of a decision depends almost entirely on how much you send. A routing check that ships a whole conversation history costs far more than one that sends the last message and a short state summary. Context discipline becomes cost discipline.
The decision layer gets its own budget line. Routing, escalation, moderation and next-step classification often make up most of an agent’s calls. Moving them to a cheaper, faster endpoint changes the economics we described in our analysis of routing by task in OpenAI’s GPT-6 guide, and makes visible which part of an agent’s spend is reasoning and which is triage.
Custody becomes a real choice. For decisions that touch regulated data, open weights you host and an API you call are different risk profiles. That difference may matter more than speed or price.
What a decision API still does not decide
A probability is not a permission. A high-confidence “approve” still passes through a deterministic policy check that you own, with thresholds in configuration and a band that goes to a person.
Calibration needs proof on your own traffic: a 0.9 should be right about nine times in ten. Early public-beta users in OpenAI’s forum are already comparing how well the formats behave, a reminder to measure before setting thresholds.
And a beta is a beta. Rate limits, pricing and behavior can change before general availability. The safest way to adopt any of these models, OpenAI’s, Cloudflare’s or your own, is behind one internal interface, so the decision layer can change vendors without the agent noticing.