← News

Cloudflare Blog launch graphic: Introducing Clef open-source decision models and RL fine-tuning platform

1 Oct 2026

Cloudflare Blog

Cloudflare open-sources Clef decision models on Workers AI

Cloudflare said it released Clef and Clef-flash, the first models trained by its Workers AI team, hosted on Workers AI and open-sourced under Apache 2.0 on Hugging Face — decision models that read a state and typed questions and return probabilities for every allowed answer, plus a new reinforcement-learning fine-tuning service.

Agents that route, block, or escalate need a probability they can act on in milliseconds. Cloudflare is putting open decision-model weights on Workers AI, on the network that already sits in front of a large share of web traffic, a bet that this hot-path classification becomes infrastructure.

On Thursday, 1 October 2026, Cloudflare said it released Clef and Clef-flash. The company blog is titled “Introducing Clef: our open-source decision models, and new RL fine-tuning platform.” The page dates the post October 1, 2026. The byline is Michelle Chen, Alex Reneau, and Kevin Flansburg. The Workers AI changelog, the same day, calls them the first models trained by the Cloudflare Workers AI team, and says they are available on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. Workers AI is Cloudflare’s service for running a model on its network. Cloudflare also put the weights on Hugging Face under the Apache 2.0 license, as Cloudflare/clef and Cloudflare/clef-flash. Weights are the numbers inside a model. Apache 2.0 is a license that lets someone use, change, and ship those weights, including in a commercial product, under the terms written in the license. Those lines are Cloudflare’s.

A decision model, as Cloudflare describes it, reads a situation and a list of typed questions, then returns a probability for every answer the caller allowed. The situation is the state. It can be a sentence, or structured data such as a record, a chat log, or the state of an application. The reply is a number for each allowed answer. There is no free-form paragraph to parse, and the changelog says there are no reasoning tokens to wait for. A token, here, is a small piece of the input the model counts. An agent, software that takes a next step, can use those probabilities to route a ticket, block a request, or hand the case to a person. Those lines are Cloudflare’s.

The changelog says one request can ask up to 64 questions, in three types. A noul question is yes or no. It returns the probability that the answer is yes. Noul is the name Cloudflare uses for that yes-or-no type. A choice question asks the model to pick one option from a set the caller writes. It returns the chosen option, a probability for each option, and a confidence value. A score question rates the state on an ordered scale the caller writes, such as no impact, minor, major, or critical. It returns a probability-weighted score and a probability for each level. The model page says the answers come back under the same question names the caller sent. Those lines are Cloudflare’s.

Cloudflare lists Clef at 27 billion parameters and Clef-flash at 9 billion. A parameter is one of the numbers that shapes how the model scores an answer. The changelog calls Clef the model for the highest-precision decisions, and Clef-flash the model for decisions that have to come back quickly, on the hot path. The hot path is the step an agent cannot skip before it acts. Both list a 64K context window. The model pages print that window as 65,536 tokens, which is 64 times 1,024. The Hugging Face card says Clef is post-trained from Qwen’s Qwen3.8-27B model. The blog says Clef-flash starts from Qwen3.5-9B. Post-trained means Cloudflare trained further on those existing models so the result scores the allowed answers directly. The changelog says Clef follows the System One API, the interface Typesafe uses for its Jev decision model, so an existing Jev integration can switch by changing the endpoint and the model name. The blog says the two models are Jev-API compatible, and that the 64K window is twice the 32K window it attributes to Jev. Those sizes and that comparison are Cloudflare’s.

Cloudflare reports speed from its own table of 43 eval benchmarks. The median, the middle result, is 209.3 milliseconds for Clef, 38.8 milliseconds for Clef-flash, and 524.1 milliseconds for Jev. A millisecond is a thousandth of a second. The changelog says that puts Clef about 2.5 times faster than Jev at the median, and Clef-flash about 13 times faster. The same table’s 95th percentile, a slow case that only about 5 percent of the results exceed, is 238.6 milliseconds for Clef, 122.4 for Clef-flash, and 536.0 for Jev. On that same blog table, a model Cloudflare calls Laya has a median of 5.8 milliseconds, and Cloudflare says Laya buys that speed by giving up quality on the quality benchmarks. These times are Cloudflare’s measurements. The posts do not cite an outside lab that re-ran the clock. Cloudflare says hosting on Workers AI adds to the speed, because the models run on graphics chips across its network, close to the user.

The changelog says that across 10 decision benchmarks, a Clef model scores highest on 7, ahead of Jev and other open decision models. That count is Cloudflare’s, from the table in the launch post. On the rows Cloudflare highlights, BFCL, a case-exact score, is 98.76 for Clef-flash and 98.47 for Clef, against 95.75 for Jev. BANKING77, a macro-F1 score, is 94.20 for Clef against 79.74 for Jev. CLINC150 with out-of-scope cases is 97.43 for Clef against 89.27 for Jev. A home-appliances task is 97.73 for Clef-flash against 52.27 for Jev. Case exact means the whole answer matched. Macro-F1 averages how well the model hits each category, so a rare category counts as much as a common one. In that same 10-row table, Jev leads When2Call and BRIGHT, and a model Cloudflare labels DiffusionGemma Jev leads PhishNChips. On Typesafe’s workflow tests, Cloudflare says a Clef model beats Jev in 3 of 4 areas: invoice processing, 64.7 for Clef against 61.8 for Jev; customer service, 76.3 for Clef and 77.0 for Clef-flash against 76.0 for Jev; and security incidents, 62.9 for Clef against 61.7 for Jev. On the fourth, agent-trace observability, Jev scores 71.6, Clef-flash 69.8, and Clef 68.5. The Hugging Face card prints a longer Decision Index table from the same internal run. Those scores are Cloudflare’s as well. The blog points readers to a live demo Cloudflare hosts. The tables are the company’s.

Clef can look at pictures. The changelog says it has a vision encoder, and contrasts that with Jev, which Cloudflare describes as text classification. A vision encoder is the part of the model that reads an image. The model pages say Clef and Clef-flash read the state as text, JSON, images, or video. JSON is a structured data format. The request schema on the Clef page documents images in particular: up to four PNG, JPEG, or WebP images, sent with the state, up to 4 megabytes and 16 megapixels each, 8 megabytes decoded in total, and a whole request of at most 13 megabytes. The page says remote web addresses for those images are not accepted. Cloudflare’s threat-intelligence example uses a tool it calls Browser Run to fetch and render a site. In that workflow, Cloudflare says Clef classified a domain in 2.2 seconds. It says gpt-oss-120b, which it calls its fastest general model, took 4.7 seconds in the same workflow and returned only two classifications. The blog’s sketch of the kind of answer is a domain that might come back as a 95 percent chance of fashion, 85 percent ecommerce, and under 1 percent phishing. That sketch is Cloudflare’s example. It does not name the domain. The 2.2 seconds and the 4.7 seconds are Cloudflare’s timings.

Cloudflare is also offering a reinforcement-learning service so a customer can fine-tune Clef for a specific job. Fine-tune means train the model further on that customer’s examples. Reinforcement learning, in this post, is that further training, aimed at decisions the customer cares about. The blog says the offer starts as hands-on work with a forward-deployed engineer team, for customers who sign up as design partners. A design partner is a customer working with Cloudflare while the service is still being shaped. The changelog links a signup for those partners. The blog says a self-serve platform comes later, after that hands-on work. The later platform, as the post describes it, would use tools Cloudflare already runs: AI Gateway, to keep a dataset of requests the customer already sends; Workers AI, to try the base Clef model; Containers, as a sandbox for scoring and replaying actions; a trainer, to update the weights; and Workers AI’s bring-your-own-model path, to put the tuned model back on the network. The post does not say that self-serve trainer is open to every customer on Thursday. Internal jobs Cloudflare names as the reason for the work: scoring Trust and Safety submissions, triaging support requests, and deciding whether a crawler is a welcome bot or an unwanted one. Those are Cloudflare’s plans for its own teams. The post does not say a customer can switch those tuned models on today.

The Clef model page lists a unit price of $0.24 per million input tokens. The Clef-flash page lists $0.09 per million input tokens. A million input tokens is the billing unit printed on those pages. The blog’s fine-tuning section does not print a dollar price for the design-partner work. The blog also says the hosted models come with a guarantee that Cloudflare does not read, store, or train on a customer’s requests or responses, unless the customer uses the fine-tuning product. That guarantee is Cloudflare’s statement.

The picture is the Cloudflare Blog launch graphic. The orange cloud mark and the word CLOUDFLARE sit at the upper left. An orange line reads CLOUDFLARE BLOG. The headline reads “Introducing Clef: our open-source decision models, and new RL fine-tuning platform.” On the right, a glass sphere holds an orange clef monogram, with small stars and thin horizontal lines behind it. The sphere sits on a white stand and a pale tilted square. An orange bar runs along the bottom. The graphic does not print a calendar date. It is the launch graphic. It is not a photograph of a data center, and it is not a chart of the benchmark table.

In plain terms, Cloudflare said on Thursday that two decision models are on Workers AI, and that the weights are on Hugging Face under Apache 2.0. Clef is the larger one, listed at 27 billion parameters. Clef-flash is the smaller one, listed at 9 billion. They take a state and typed questions and return probabilities, so software can route, block, or escalate. Cloudflare’s own medians are 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash, against 524.1 milliseconds for Jev. Cloudflare says a Clef model leads 7 of the 10 decision benchmarks in its launch table. A request can include up to four images. Fine-tuning is open to design partners, and Cloudflare describes a self-serve version as a later step. The speeds, the benchmark scores, and the domain-classification example are Cloudflare’s.

RELATED

ONLINE…

Comments

guidelines

Loading…

Loading…

Sources