Cloudflare has launched the Clef-omni model, an open-weight decision model capable of analyzing text, images, audio, and video through a single processing pipeline. The announcement follows the launch of Clef and Clef-flash last week and targets use cases requiring decisions constrained by a specific schema, rather than relying on a chain of separate models to convert audio to text or separate visual channels from video.
One Model for Multimodal Inputs
Clef-omni supports audio files in WAV and MP3 formats and video in MP4 and WebM formats, in addition to text and images. Cloudflare presents an example involving the inspection of a device using an image of the product unit, an audio recording while it is operating, and a video of the fan, followed by structured questions about whether the model number is visible, whether the sound is normal, and whether the fan is working.
The model is built on Qwen3-Omni-30B-A3B-Instruct using a Mixture of Experts (MoE) architecture, with the core understanding component used while the text-to-speech components are omitted. Rather than generating output tokens like large language models, Clef directly calculates scores for candidate options within a defined schema, using a two-stage attention-routing mechanism and embedded grammar rules to preserve the meaning of the options.
Cloudflare says that text decisions are returned with a median latency of approximately 130 milliseconds, while image inputs take about 150 milliseconds and audio inputs require a few hundred milliseconds. A 21-second video with audio is evaluated in about 1.5 seconds through a single API request. The company has made Clef-omni’s weights available on Hugging Face, along with developer documentation.
What Changes in Practice?
The new design reduces the need to build sequential processing pipelines to transcribe speech or break down video before making a decision. This suits inspection and classification tasks requiring structured answers, such as invoice processing, customer service, and security incident monitoring, provided that the model’s quality is sufficient for the specific domain.
However, the test results do not outperform other models on every task. Clef-omni scored 98.2% on the BFCL test for matching cases, 92.7% on API-Bank, and 94.8% on BANKING77, but it fell below Clef and Clef-flash in some tests, such as Home Devices, When2Call, and PhishNChips. Therefore, the addition of multimodal capabilities does not mean overall superiority in every scenario.
Clef-flash Price Cut and Clef Speed Boost
Cloudflare reduced Clef-flash’s price from $0.09 to $0.038 per million input tokens, while Clef’s price remains $0.24, and Clef-omni launched at $0.15 per million input tokens. At the same time, the context window in the hosted version of Clef-flash was reduced from 64,000 to 24,000 tokens. The company says that only 0.24% of requests exceed the new limit, while the Hugging Face weights remain trained to support a window of up to 256,000 tokens when self-hosted.
Cloudflare also improved the speed of hosted Clef on Workers AI. Median latency fell from 262 to 152 milliseconds for inputs of approximately 800 tokens, and from 616 to 305 milliseconds for inputs of approximately 3,400 tokens, representing a speedup of up to 2.0 times. The company attributes part of the improvement to the use of SGLang, with contributions set to appear in version 0.5.22.
Why Does This Announcement Matter?
The announcement presents a practical direction for decision models: processing different inputs and extracting structured options instead of producing long text. Teams within Cloudflare have used these models to detect spam messages in GitHub, inspect EmDash system extensions for phishing, monitor personal metadata, and detect malicious domains. However, these examples are internal, and the mixed test results and context-window limitations in Clef-flash remain factors that should be evaluated before adopting the model in sensitive operations.