Fetching latest headlines…

Dev

Decision Models Belong on the Agent Hot Path: What Cloudflare Clef Changes

Dev.toUnited States · NORTH AMERICA

Cloudflare released Clef and Clef-flash on October 1, 2026, and the useful part is not that another model family exists. The useful part is that these models are designed for a narrow job that large l...

0 views0 likes0 comments

Cloudflare released Clef and Clef-flash on October 1, 2026, and the useful part is not that another model family exists. The useful part is that these models are designed for a narrow job that large language models often perform badly: choosing among a bounded set of actions quickly, with typed outputs and calibrated probabilities.

That distinction matters for agent systems. Many agent loops use a large language model for everything: interpret an event, decide whether it is risky, choose a route, pick a tool, select a priority, and only then generate text or execute an action. The result is flexible, but it is also expensive, variable, and difficult to test.

Clef points toward a different architecture. Put a small decision model on the hot path for bounded choices, then call a larger model only when the task actually needs open-ended reasoning or generation.

Cloudflare says Clef is based on a frozen Qwen3.8-27B backbone, while Clef-flash uses a frozen Qwen3.5-9B backbone. The company trains routing heads and low-rank adapters rather than turning the whole system into another general text generator. The output is a probability distribution over allowed schema choices.

That sounds modest. It changes a surprising amount.

A Decision Is Not the Same Thing as a Generation

Consider a support agent receiving a customer message. The system may need to answer three questions before anything else happens:

  • Is the request urgent?
  • Which team should receive it?
  • Should the system escalate to a human?

A language model can answer those questions, but the task itself is bounded. There is no need to generate paragraphs, chain through tool calls, or produce prose that another parser must interpret.

A decision model can represent the same job as typed choices:

schema = {
    "urgent": [True, False],
    "team": ["billing", "security", "support"],
    "escalate": [True, False],
}

The useful output is not a sentence saying that a ticket appears urgent. The useful output is a stable record with probabilities attached to each allowed choice.

That gives the caller something operational rather than conversational.

The distinction becomes more important as agents move into production systems. A free-form model may decide that a ticket is "probably urgent" or invent a new category such as "priority support." A schema-bound model cannot create a category that the downstream system does not understand.

This is less expressive by design. For routing, policy gates, moderation decisions, retries, model selection, and workflow transitions, that is often a feature.

Clef Treats Schema Choices as the Product

Cloudflare describes Clef as using a two-stage attention routing process. Each valid choice extracts context from the prompt, fields can cross-attend with other fields, and the system scores only the options that the caller declared.

The important engineering idea is not the exact attention layout. It is that the schema is part of inference rather than an instruction the model may or may not follow.

A typical agent router could ask for several fields at once:

{
  "input": "Production API latency doubled after a deploy",
  "questions": {
    "severity": ["low", "medium", "high", "critical"],
    "owner": ["application", "database", "network"],
    "rollback": ["yes", "no"]
  }
}

The caller already knows every legal answer. The model's job is to rank them.

That is a cleaner contract than asking a text model to emit JSON and then validating, repairing, or retrying the result. Structured generation has improved, but it still solves a broader problem than many control-plane decisions require.

This is why decision models fit naturally beside large language models rather than replacing them. The decision model narrows the path. The language model handles the parts that remain open-ended.

The Latency Numbers Change Where a Model Can Sit

Cloudflare's published benchmark table reports a median latency of about 209.3 milliseconds for Clef, 38.8 milliseconds for Clef-flash, and 524.1 milliseconds for Jev across its benchmark runs. Its p95 figures are about 238.6 milliseconds for Clef, 122.4 milliseconds for Clef-flash, and 536 milliseconds for Jev.

Those measurements are Cloudflare's own benchmark results, so they should not be treated as universal production numbers. Hardware, region, payload size, and schema complexity all matter.

The architectural implication is still useful.

A model that consistently returns a bounded routing decision in tens of milliseconds can sit in places where a general language model feels too heavy. Examples include request admission, abuse scoring, tool routing, fallback selection, alert triage, queue assignment, and policy checks before an expensive model call.

A simple router might look like this:

def handle_request(request):
    decision = clef_decide(request)

    if decision["risk"] == "high":
        return escalate(request)

    if decision["needs_reasoning"] == "yes":
        return call_large_model(request)

    return run_fast_path(request, decision)

The model is no longer the whole application. It is one bounded component in a larger system.

That makes latency easier to budget and failure easier to isolate.

Probability Calibration Is More Useful Than Confident Prose

One of the more interesting training details is Cloudflare's use of a Brier loss for probability calibration. The company also describes a reinforcement-learning objective called Reinforcement Learning for Calibrated Decisions, or RLCD.

Calibration matters because production systems rarely need only a label. They need to know how much confidence to place in that label.

Suppose a moderation gate returns the same top answer for two requests:

request A: allow 0.98, review 0.02
request B: allow 0.54, review 0.46

Treating those as identical "allow" decisions would throw away useful information.

A production controller can set explicit thresholds. High-confidence choices can continue automatically. Borderline choices can go to a larger model or a human. Very low-confidence cases can fail closed.

This is a better fit for policy systems than trying to infer confidence from the tone of generated text. Fluent prose is not a calibrated score.

The same pattern applies to incident routing, fraud detection, email classification, customer support queues, and tool selection. The more expensive the wrong action is, the more valuable calibrated uncertainty becomes.

Frozen Backbones Make the Fine-Tuning Story More Practical

Cloudflare says it freezes the Qwen backbones and trains routing components plus rank-256 low-rank adapters. That choice matters because it reduces the amount of model state that must change for a specialized task.

The company is pairing Clef with a new reinforcement-learning service. Its proposed pipeline uses existing Cloudflare infrastructure: AI Gateway to collect traffic, Workers AI for rollouts, Containers as scoring sandboxes, a Trainer component to update weights, and Workers AI with bring-your-own-model support for redeployment.

The loop can be expressed as a conventional control system:

production traffic
    -> capture examples
    -> score decisions
    -> train adapters
    -> evaluate calibration
    -> deploy candidate
    -> compare with current model

That is more operationally interesting than "fine-tune a model" as a standalone feature.

The hard part of applied reinforcement learning is often not the optimizer. It is constructing repeatable environments, capturing representative data, defining rewards, replaying failures, and deciding when a new model is safe enough to serve.

Cloudflare already operates traffic, compute, sandboxing, and model serving layers. Its RL product is an attempt to connect those pieces into one feedback loop.

Whether that service becomes broadly useful will depend on tooling quality and evaluation discipline, not just training speed.

Decision Models Can Cut Large Models Out of the Boring Branches

A typical agent architecture has many decisions that do not require language generation.

Which model should handle this request?

Should the system retry or stop?

Which tool is eligible?

Is the user asking for an action or an explanation?

Does this event belong in the security queue or the operations queue?

Should an expensive retrieval step run at all?

Large models can answer all of these. That does not mean they should.

If a system calls a large model for every branch, the cost is not only inference. It also inherits larger variance, longer tail latency, more prompt surface, more formatting checks, and more opportunities for prompt injection to influence control logic.

A bounded decision layer shrinks that attack and validation surface. The application declares legal outputs first. The model assigns probability to them second.

That does not remove security concerns. Inputs can still be adversarial, and a classifier can still be wrong. It does make the contract easier to reason about.

For many agent systems, the best design may be a cascade:

rules for obvious cases
    -> decision model for bounded ambiguity
    -> large model for open-ended reasoning
    -> human for high-impact uncertainty

Each layer handles the cases that justify its cost and flexibility.

Open Weights Matter Because Routing Logic Is Infrastructure

Cloudflare released Clef and Clef-flash under the Apache 2.0 license and says the hosted versions are compatible with the Jev API.

That is important because routing logic is not merely another content feature. Once a decision model sits in the request path, it becomes infrastructure.

Infrastructure benefits from portability. Teams may want to run the same decision model locally, inside a private network, near a data source, or on a different inference stack. Open weights make it possible to inspect behavior, benchmark on internal data, and move serving locations without redesigning the application contract.

Compatibility also reduces switching cost. If multiple decision models accept the same bounded schema interface, applications can compare them without rewriting every caller.

That creates a more interesting competitive dimension than raw benchmark rank. A team can ask which model is best calibrated on its own traffic, which one meets latency targets, and which one behaves most predictably under distribution shift.

Those are infrastructure questions, not chatbot questions.

The Main Risk Is Treating Every Choice as a Classification Problem

The appeal of fast structured decisions can produce the opposite mistake: forcing open-ended reasoning into a schema that is too small.

Some tasks genuinely need exploration. A novel production incident may not belong to any known category. A code-review problem may require tracing state across several files. A user may ask something that changes the available choices themselves.

A decision model is strongest when the action space is already known.

That suggests a useful design test. Before adding a decision model, ask four questions:

  1. Are the legal outputs known before inference?
  2. Can a wrong choice be detected or escalated?
  3. Is latency or cost important enough to justify a specialized path?
  4. Does the system benefit from explicit probabilities rather than prose?

If the answer to those questions is mostly yes, a decision model is a strong candidate.

If the model needs to invent the action space, explain tradeoffs, synthesize evidence, or create a plan that did not exist before the request arrived, a general reasoning model remains the better tool.

The Better Agent Architecture Is More Heterogeneous

Clef is interesting because it argues against the idea that every agent improvement requires a larger general model.

Production systems already use different components for storage, search, queues, policy, and execution. Model architecture is starting to follow the same pattern.

One model may extract. Another may rank. Another may make a bounded decision. Another may reason across a large context. A deterministic program may perform the final write.

That looks less magical than a single autonomous model controlling everything. It is also easier to measure.

The practical lesson is not that decision models will replace language models. It is that many agent systems are wasting language-model capacity on decisions that should be explicit, typed, and cheap.

Cloudflare's Clef release makes that separation concrete. The most useful question for an agent designer is now simpler: which parts of the workflow truly require generation, and which parts only require a reliable choice among known options?

Sources

Originally published on Dispatch.

Comments (0)

Sign in to join the discussion

Be the first to comment!