Cloudflare has expanded its Clef family of open-weight decision models with Clef-omni, a model that natively takes in audio, video, images and text in a single pipeline - and scores structured decisions across all of them without generating a word of output. The company announced the release on October 9 alongside open weights on Hugging Face, a price cut for Clef-flash, and inference speedups of up to 2x for the existing Clef.
The modality expansion is the headline. Since TypeSafe's Jev debuted the decision-model category, such models have been largely text-only; Cloudflare's original Clef supported images and arrays of video frames, which forced developers to split audio and video into preprocessed pieces. Clef-omni accepts WAV or MP3 audio and MP4 or WebM video directly, so a developer can ask one model whether a machine sounds like it is running normally, whether a fan is spinning, or whether a video frame shows a defect - without standing up a speech-to-text transcription stage or a frame extractor first.
Under the hood, Cloudflare built the model on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts backbone, keeping the comprehension stack while discarding the base model's text-to-speech output components. In production, Clef-omni runs a quick prefill pass across the full payload, scoring all modalities and valid parameter options simultaneously; media elements map into a unified sequence where video and audio stay synced for joint processing. Candidate values are pulled from internal embeddings via a two-stage attention routing scheme, with a built-in lexical grammar preserving option semantics. Training followed the same recipe as the original Clef: a frozen Qwen3 backbone, low-rank adapters, and post-training that combines label-smoothed cross-entropy loss with Brier score calibration.
The performance figures are company-reported. Cloudflare says median response time is about 130 milliseconds for text-only decisions and 150 milliseconds for images, with audio clips taking a few hundred milliseconds and a 21-second video with sound about 1.5 seconds. On its published evaluations, Clef-omni posts a 69.38 exact-match rate on home-appliance control decisions and 73.27 accuracy on anti-phishing classification, trading wins and losses with the earlier Clef and Clef-flash on tool-calling and API-calling tasks. No independent benchmark has covered the family yet.
The pricing move sharpens the competitive picture: Clef-flash now costs $0.038 per million input tokens, down from $0.09 - roughly a 58 percent cut - which Cloudflare says makes it cheaper than Jev, the model that started the category. The hosted version's context window has been trimmed from 64k to 24k tokens. TypeSafe, for its part, closed a roughly $870 million round at a $7.5 billion valuation last week, so the decision-model niche now has both a valuation benchmark and a price war.
Cloudflare's account of how fast this happened is almost as notable as the model: the company says the original Clef went from a Friday-evening decision to a Thursday launch in under a week, training over a weekend. Whatever one thinks of week-long model sprints, the direction is clear - the arbitration layer that sits in front of expensive LLMs is getting multimodal, cheaper, and faster, and it is being contested by infrastructure companies rather than labs.
Comments (0)
Log in to join the discussion
Log InNo comments yet