Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

Cloudflare Open-Sources Clef, a 27B Decision Model That Claims to Beat Jev on Speed - With Apache 2.0 Weights

Cloudflare Open-Sources Clef, a 27B Decision Model That Claims to Beat Jev on Speed - With Apache 2.0 Weights

Cloudflare has released Clef and Clef-flash, its first in-house decision models, with Apache 2.0 weights on Hugging Face and hosting on Workers AI. The company reports median latencies of 209.3ms and 38.8ms against Jev's 524.1ms, and says Clef tops 7 of 10 decision benchmarks. All benchmark numbers are self-reported. Clef is API-compatible with TypeSafe's Jev, making it a drop-in replacement.

Cloudflare has entered the decision-model race with both models at once. Clef and Clef-flash, the first models trained by the company's Workers AI team, answer structured questions with probabilities instead of generating text - and unlike the two competing decision models released in the previous two days, Cloudflare is giving away the full weights under Apache 2.0 on Hugging Face.

The timing is the story. TypeSafe's Jev turned "decision models" - fast, bounded, machine-readable answers for agent loops - into a product category in September. OpenAI followed with its Luna-based Decisions API at DevDay on September 30, and Amazon shipped Strands Decider 2B on October 1. Clef arrives the same week as the third entry, and the only one of the three shipping open weights at scale: a 27.4-billion-parameter model built on a frozen Qwen3.8-27B backbone with a trained routing head and rank-256 LoRA adapters, plus a 9-billion-parameter Clef-flash built on Qwen3.5-9B.

The architecture skips autoregressive generation entirely. Clef reads the input state and all questions in a single prefill pass, then scores every valid answer option in parallel - up to 64 questions per request, each typed as noul (a yes/no probability), choice (one option from a set), or score (a rating against an ordered rubric). Context runs to 64,000 tokens, and each request can carry up to four images alongside text and JSON, something Jev, which is text-only with a 32k window per request, does not support.

On Cloudflare's own benchmark runs, the speed claims are aggressive: across 43 runs, median latency of 209.3ms for Clef and 38.8ms for Clef-flash against 524.1ms for Jev - 2.5x and 13x faster respectively, with p95s of 238.6ms, 122.4ms and 536.0ms. On accuracy, Cloudflare says Clef scores highest on 7 of 10 decision benchmarks, including BANKING77 macro-F1 of 94.20 versus Jev's 79.74, and BFCL case-exact of 98.47 versus 95.75. On TypeSafe's own workflow evaluations, Clef says it beats Jev in three of four areas - invoice processing, customer service and security incidents - while losing on agent trace observability. All of these figures are self-reported and have not yet been reproduced on the public Decision Index.

The trade-off is price and openness, in both directions. Clef costs $0.24 per million input tokens and Clef-flash $0.09 on Workers AI, versus Jev's $0.042 - roughly six times the price - but a developer can instead download the weights and run them anywhere, given roughly 85GB of VRAM for Clef or 41GB for Clef-flash at single concurrency with a 64k context. And there is a nuance worth noting: Cloudflare describes the release as open source, but product manager Michelle Chen confirmed to The Register that the training datasets are not public - open weights, not fully open source.

The API is fully Jev-compatible, following the System One request shape, so existing integrations can switch by changing the endpoint. Cloudflare is also launching a reinforcement-learning fine-tuning service built on RLCD (Reinforcement Learning for Calibrated Decisions), initially as a hands-on engagement with its forward-deployed engineers, with a self-serve version planned for AI Gateway and Containers. Internally, Cloudflare's threat intelligence team already uses Clef to classify website domains - 2.2 seconds per domain against 4.7 seconds for gpt-oss-120b.

Three decision models in two days means the category is no longer a product; it is becoming a wire format. For developers, that is unambiguously good - routing, triage and guardrail logic can now be written once and pointed at any of three vendors, or self-hosted outright. For the vendors, the interesting question is whether decision models go the way of commodity infrastructure, where open weights win on trust even when the hosted alternative is faster. Cloudflare is betting they do.

Comments (0)

Log in to join the discussion

Log In

No comments yet