THE SIGNAL IN ONE SENTENCE

Sometimes an AI system needs to write a paragraph. Sometimes it needs to decide whether a support ticket goes to billing, fraud or a person before the customer finishes making coffee. Cloudflare has built two models for the second job. The company released Clef and Clef-flash on October 1. Instead of producing free-form text one token at a time, the models receive a state and a set of typed questions with allowed answers. They return probabilities for those answers in one pass. No essay. No polite preamble. No parser rummaging through prose for the actual choice. The models are available through Workers AI, Cloudflare's hosted inference platform. Cloudflare also published the weights, model cards and supporting code under the Apache 2.0 license. Clef is a 27-billion-parameter multimodal model based on Qwen3.8-27B. Clef-flash is the smaller 9-billion-parameter version. Both support a context window of up to 64,000 tokens, according to Cloudflare. That is the first half of the release, and it is genuinely portable. A team can download the models, inspect the code and run them on infrastructure it controls. The second half points in the opposite direction. Cloudflare also announced a reinforcement-learning fine-tuning service. The initial version is a hands-on engagement with the company's forward-deployed engineering team. Cloudflare says it plans a self-service platform that will capture a customer's data, fine-tune a decision model and redeploy the result, all on Cloudflare. The plain signal is that Cloudflare has separated a useful model from an even more valuable loop. The weights are open. The managed system for collecting outcomes, improving the model and putting the new version back into production is being built as a Cloudflare product. That distinction matters because the loop can become harder to move than the model. A decision model is meant for bounded actions. Ask whether a transaction should be approved, reviewed or rejected. Ask which tool an agent should call. Ask which of five queues should receive a ticket. Ask whether an image matches a policy category. Each answer is constrained to a defined option, and the model returns a score or probability for each allowed choice. This structure removes a familiar source of agent trouble. A generative model might answer, "I would probably send this to billing, although fraud may want a look." Software then has to parse that sentence, decide whether "probably" is decisive and hope the model did not invent a fourth queue. A decision model can return billing at 0.81, fraud at 0.14 and human review at 0.05. The numbers are easier for software to consume. They are not automatically true. Cloudflare's architecture is built for speed. The company says Clef uses a Qwen backbone for a prefill-only pass, then scores schema choices in parallel through a specialized routing head. It does not generate a chain of text. Clef and Clef-flash use frozen Qwen backbones with low-rank adapters and a decision head, according to the technical description. Cloudflare reports that this design made the models faster than Jev, a prior open decision model used as a comparison. Across 43 benchmark runs, the company measured median latency of 209.3 milliseconds for Clef, 38.8 milliseconds for Clef-flash and 524.1 milliseconds for Jev. It says Clef beat Jev on three of four typesafe workflow evaluations, while Jev won the agent-trace observability task. The benchmark table is more interesting than a single winner label. Clef led Jev on a banking-intent test and a phishing classification test. Clef-flash led on a home-appliance tool-use test. Jev performed better on When2Call and the BRIGHT retrieval benchmark. On PhishNChips, another comparison model called DiffusionGemma Jev beat every Clef variant. That mixed result is healthy evidence. It says different decision tasks reward different training choices. It also means "fast decision model" should not be translated into "best model for every decision." All of these benchmark and latency results come from Cloudflare. The company chose the tests, ran the systems and published the measurements. No independent replication is cited in the release. The announcement also names no launch customer running the reinforcement-learning service in production. Cloudflare does describe an internal use case. Its threat-intelligence workflow can fetch and render a web page, then classify its domain. The company reports that Clef completed the full workflow in 2.2 seconds, compared with 4.7 seconds for gpt-oss-120b. That is useful directional evidence for a constrained classifier near an edge network. It is not a neutral measurement of every possible deployment. The open-weight release gives independent teams a way to test the claim. Cloudflare's model card says Clef is distributed as standard sharded safetensors with the schema head and code needed to run it. The documented setup was tested with PyTorch 2.11 and Transformers 5.10.2 on a single Nvidia H200. That is portable in the licensing sense. It is not a promise that a 27-billion-parameter model will be cheap to serve on ordinary hardware. Clef-flash is likely the more practical starting point for teams that care about latency and cost. Smaller does not mean harmless, however. A fast wrong decision in a hot path can produce damage with impressive efficiency. Probabilities need calibration. If a model assigns 0.9 to one hundred similar cases, roughly ninety should be correct under the conditions for which the system was calibrated. A score that looks precise can drift when users, products, fraud patterns or policies change. A team should test whether 0.9 still behaves like 0.9 on its own data, not merely whether the top answer often wins. Thresholds also encode values. A company may auto-route a low-risk support ticket at 0.7 but require 0.98 before blocking an account. False positives and false negatives do not cost the same thing. The right threshold depends on who absorbs the error, whether the action is reversible and how quickly a human can intervene. This is where the reinforcement-learning loop becomes attractive. Record the decision, observe the outcome, label what worked, adjust the model and redeploy. Cloudflare says its planned platform will connect these steps. It also says current fine-tuning starts with its engineering team and will move toward a self-service system. The release does not establish that the self-service platform exists today. Nor does it describe every governance control that customers will need around consent, retention, deletion, label quality, rollback and cross-version audits. Those details are not side quests. Outcome data can be more sensitive than the original prompt. A support label may reveal a customer's financial trouble. A fraud disposition may contain investigative logic. A moderation result may encode an organization's policy judgments. Capturing that material for training needs explicit boundaries and a durable record of which data changed which model. Cloudflare says ordinary Workers AI requests and responses are not read, stored or used for training unless a customer uses fine-tuning. That is an important baseline. Fine-tuning is precisely the case in which collection becomes the feature, so teams need a separate data agreement and technical controls for that workflow. Open weights reduce one form of lock-in. A company can keep a copy, run local evaluations and build an exit path. A managed outcome loop can create another form. If labels, evaluation sets, deployment history and rollback tools live only in one platform, moving the model later may not recreate the system that made it useful. The practical answer is not to reject the managed service. It is to keep the learning record portable. Export the training examples in a documented format. Version the decision schema. Preserve the evaluation set and threshold history. Record which model, adapter and policy produced each consequential action. Keep a holdout set outside the improvement loop. Test a local copy before depending on the hosted one. Define what must happen if Cloudflare is unavailable or a new fine-tune performs worse. Start in shadow mode. Let the model score real cases without controlling them. Compare its choices with existing decisions and eventual outcomes. Break performance down by language, customer group and unusual case type. A single average can hide a terrible minority experience. Then add selective automation. Send obvious, reversible cases down the fast path. Route uncertain, novel or high-cost cases to a person. An abstention path is not a failure. It is one of the most useful answers a decision system can produce. Good applications are easy to imagine. Ticket routing, tool selection, abuse triage, invoice classification, incident severity, product matching and domain categorization all contain bounded choices. Multimodal input adds cases in which the decision depends on an image or video as well as text. The release also suggests a larger shift in agent design. Many agent systems currently use one generative model for planning, writing, classification, routing and permission choices. That is convenient, but it turns every small fork into a miniature conversation. A specialized decision model can occupy the hot path while a larger generative model handles the parts that actually need language. This architecture may be faster, cheaper and easier to test. It is also easier to over-trust because the output looks clean. Three probabilities fit neatly into a database. A paragraph at least advertises its ambiguity. Cloudflare has made the bounded choice a first-class AI product. The open weights let developers verify whether the model fits their work. The hosted service lets them try it near Cloudflare's edge. The planned learning loop could make the model increasingly specific to one operation. The bargain is straightforward. Use the speed, but keep a human path. Use the probabilities, but measure calibration. Use the open weights, but also export the evidence that makes a tuned model valuable. The model can leave the cloud. Make sure the learning can leave with it.

01

WHAT ACTUALLY CHANGED

Cloudflare released Clef and Clef-flash on October 1, 2026.

Both models answer typed, bounded questions with probabilities instead of generating free-form text.

Clef is a 27-billion-parameter multimodal model and Clef-flash is a 9-billion-parameter version.

Cloudflare says both models support context windows up to 64,000 tokens.

The models are available through Workers AI at Cloudflare's edge.

Cloudflare published the model weights, model cards and supporting code under the Apache 2.0 license.

The Clef API is compatible with Jev and supports up to 64 questions in one request.

Cloudflare introduced a reinforcement-learning fine-tuning service that begins with hands-on engineering support.

The company plans a self-service system for data capture, fine-tuning and redeployment on Cloudflare.

The current release does not establish independent benchmark replication or name production launch customers for the fine-tuning service.

02

WHY THIS MATTERS

A bounded answer can remove output parsing and reduce the chance that an agent invents an unsupported action.

Parallel option scoring can make small routing and classification decisions much faster than text generation.

Open weights let teams test the model locally, inspect its serving code and retain an exit path.

A managed training loop can become a stronger source of platform dependence than the model weights themselves.

Probability outputs make thresholds explicit, but they still require calibration on current, local data.

Fast decisions can amplify errors when they directly block accounts, route money or enforce policy.

Outcome data used for reinforcement learning may contain sensitive operational and customer information.

Separating decision models from generative models can make agent systems cheaper, more testable and easier to govern.

Portable labels, evaluation sets and version history are necessary if open weights are to provide meaningful independence.

FIG. 285PUT A BOUNDED MODEL IN THE HOT PATH WITHOUT MAKING IT THE BOSS
1NAME THE DECISION→
2DEFINE ALLOWED ANSWERS→
3SET THE COST OF ERRORS→
4RUN A SHADOW TEST→
5CALIBRATE THRESHOLDS→
6ESCALATE UNCERTAINTY→
7LOG OUTCOMES→
8RETRAIN WITH CONSENT→
9REDEPLOY OR ROLLBACK
The fast path should handle clear, reversible choices. Uncertain and costly cases still need a person, a record and a way back.

03

WHERE IT COULD HELP

  • Route support requests among billing, technical, fraud and human-review queues.
  • Choose which approved tool an agent should call next.
  • Classify domains, messages or images for an abuse-triage queue.
  • Assign incident severity before waking an on-call team.
  • Match a product request to a bounded catalog or shortlist.
  • Score invoices or documents for straight-through processing versus review.
  • Run a model in shadow mode beside an existing workflow before allowing automation.
  • Set different thresholds for reversible and irreversible actions.
  • Export examples, labels, evaluation sets and model versions before each fine-tuning run.
  • Send uncertain, novel and high-cost cases to a human instead of forcing a choice.

KEEP A HAND ON THE WHEEL

Cloudflare's latency, benchmark and workflow results were produced by Cloudflare and have not been independently replicated in the cited material. Performance varies by task, and competing models win some published comparisons. The model cards document a tested H200 setup, not the cost or speed of every local deployment. The reinforcement-learning service currently begins with Cloudflare's engineering team, while the self-service capture, fine-tune and redeploy platform is planned. No production customer, independent security review or complete governance specification is named. Probability outputs are not guarantees; teams must test calibration, error costs, subgroup performance and drift on their own data.

04

TERMS WORTH KEEPING

SOURCES AND VERIFICATION STATUS

This article was written from the materials below. Product claims and dates were checked against those sources on October 2, 2026.

PUBLICATION RECEIPT: Original publication. Facts checked immediately before publication against Cloudflare's October 1 release, Workers AI changelog and both Apache 2.0 model cards.

THE PUBLICATION ENGINE

WANT A SIGNAL OF YOUR OWN?

We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.

WORK WITH US