← All insights

News analysisLens: United States4 min read

Your agent should not ask a frontier LLM every small question. Cloudflare's Clef argues for decision models

Cloudflare released Clef and Clef-flash, open models that return probabilities over allowed answers instead of text. They point to an agent architecture with a separate, cheaper decision layer.

Listen to this article · 6 min

AI-generated narration of the full article.

A silhouetted signpost with two arrows at sunset, in La Madre duotone, beside the words Decide, not write
Photo: rawpixel (CC0)

Inside a typical agent loop, a large model is asked many small questions: is this message urgent, which team owns it, is this request allowed, should a person look at it. Each answer comes back as text that the application then parses. It works, but it is slow, expensive and hard to measure.

On October 1, Cloudflare released a different kind of model for exactly those questions.

What Cloudflare released

Clef and Clef-flash are what Cloudflare calls decision models. They read an input state plus a set of typed questions and return a probability for every allowed answer, instead of generating free-form text. Cloudflare supports three question types: yes or no, choose one of a list, and a score against a rubric.

  • Clef has 27 billion parameters; Clef-flash has 9 billion. Both take 64K tokens of context, accept text and up to four images, and are released under Apache 2.0 with weights on Hugging Face.
  • Both are generally available on Workers AI.
  • Cloudflare names support triage, threat intelligence and website classification, trust and safety scoring, agent guardrails and visual classification as uses.
  • Fine-tuning with reinforcement learning starts as a hands-on design-partner program; Cloudflare says self-serve comes later.

Cloudflare reports that Clef-flash had a median latency of 38.8 milliseconds in its comparison against another decision model, roughly 13 times faster, and that Clef led 7 of 10 decision benchmarks. These are Cloudflare’s own measurements; the launch material does not name an independent evaluator, and you should test on your own data before relying on them. Cloudflare also notes the obvious trade-off: Clef-flash gives up some accuracy for speed. Neither the changelog nor the launch post we reviewed lists pricing, so check the Workers AI pricing page.

Why a decision layer matters

The interesting part is not the speed number. It is the shape of the output.

A probability can carry a policy. When a model returns “escalate: 0.83”, you can decide that anything under 0.9 goes to a person and anything over 0.97 proceeds. That threshold is a business decision written in configuration, not a hope that the prompt is followed.

Typed answers are measurable. Allowed answers make evaluation a classification problem, with precision, recall and confusion matrices that risk and audit teams already understand. Free text routing needs a judge to grade it.

Small, fast models fit the hot path. Routing and screening happen on every request. Moving them off the frontier model leaves that model for the steps that need reasoning, which is the cost logic we described in our analysis of routing by task in OpenAI’s GPT-6 guide.

Open weights keep options open. Apache 2.0 weights mean the same decision model can run outside Cloudflare if custody or latency demands it.

A compound agent with a decision layerGENERATIVEPROBABILISTIC, TYPEDDETERMINISTIC01Retrievecontext02Reason andplan (largemodel)03Decide: route,block,approve,escalate04Deterministicpolicy check05Execute toolor hand to apersonOpen-ended reasoningProbabilities over allowed answers
  1. Retrieve context
  2. Reason and plan (large model)
  3. Decide: route, block, approve, escalate
  4. Deterministic policy check
  5. Execute tool or hand to a person
  • Generative: Retrieve context · Reason and plan (large model)
  • Probabilistic, typed: Decide: route, block, approve, escalate
  • Deterministic: Deterministic policy check · Execute tool or hand to a person

Open-ended reasoningProbabilities over allowed answers

The decision model answers narrow questions fast. The policy layer, not the model, decides what is permitted.

What a decision model does not replace

A probability is not a permission. If the decision model says “approve” with high confidence, a deterministic policy layer still checks whether this user, this agent and this action are allowed. The model narrows the choices; the rules decide. We made the same point about keeping deterministic steps deterministic in Snowflake’s accounts payable system.

Calibration also needs proof. A model that says 0.9 should be right about nine times out of ten on your traffic. Measure that before you set thresholds, and again after any fine-tuning or model update.

And some decisions carry legal weight. If a decision model screens consumer applications in the U.S., a decline that touches credit falls under the Equal Credit Opportunity Act, and Regulation B requires the lender to give specific reasons for adverse action. A probability score is not a reason; the system has to record why the outcome happened in terms a customer can be told.

What to do now

  1. List the small decisions in your agents: routing, escalation, moderation, tool choice, next-step classification.
  2. Estimate their share of model calls and cost. In many agents they are the majority of calls.
  3. Pilot a decision model on one of them with a labeled set from your own history, and compare against your current LLM prompt on accuracy, latency and cost.
  4. Set thresholds as configuration, with an explicit band that goes to a person.
  5. Keep the policy check separate and log both the probability and the rule that allowed or blocked the action.
  6. Watch calibration over time, as you would for any classifier in production.

The bottom line

Clef is one product, but the pattern is larger: agents are becoming compound systems where different models answer different kinds of questions. Give the narrow, frequent decisions to a model built to decide, keep the reasoning model for reasoning, and let deterministic policy have the last word.

Have an AI use case stuck between prototype and production?

Tell us what you’re trying to ship. We’ll reply with honest next steps.

Discuss a use case