Artificial intelligence

Cloudflare Launches Clef-omni to Process Text, Images, Audio, and Video Through a Single Model

Cloudflare has added the Clef-omni model to its open-weight decision model family, with native support for text, images, audio, and video within a single API request. It also cut the price of Clef-flash and increased Clef’s speed by up to two times, while reducing Clef-flash’s hosted context window from 64,000 to 24,000 tokens.

2026-10-09
4 min read
4 views
certi.news Editorial Team
Cloudflare Launches Clef-omni to Process Text, Images, Audio, and Video Through a Single Model

Cloudflare has launched the Clef-omni model, an open-weight decision model capable of analyzing text, images, audio, and video through a single processing pipeline. The announcement follows the launch of Clef and Clef-flash last week and targets use cases requiring decisions constrained by a specific schema, rather than relying on a chain of separate models to convert audio to text or separate visual channels from video.

One Model for Multimodal Inputs

Clef-omni supports audio files in WAV and MP3 formats and video in MP4 and WebM formats, in addition to text and images. Cloudflare presents an example involving the inspection of a device using an image of the product unit, an audio recording while it is operating, and a video of the fan, followed by structured questions about whether the model number is visible, whether the sound is normal, and whether the fan is working.

The model is built on Qwen3-Omni-30B-A3B-Instruct using a Mixture of Experts (MoE) architecture, with the core understanding component used while the text-to-speech components are omitted. Rather than generating output tokens like large language models, Clef directly calculates scores for candidate options within a defined schema, using a two-stage attention-routing mechanism and embedded grammar rules to preserve the meaning of the options.

Cloudflare says that text decisions are returned with a median latency of approximately 130 milliseconds, while image inputs take about 150 milliseconds and audio inputs require a few hundred milliseconds. A 21-second video with audio is evaluated in about 1.5 seconds through a single API request. The company has made Clef-omni’s weights available on Hugging Face, along with developer documentation.

What Changes in Practice?

The new design reduces the need to build sequential processing pipelines to transcribe speech or break down video before making a decision. This suits inspection and classification tasks requiring structured answers, such as invoice processing, customer service, and security incident monitoring, provided that the model’s quality is sufficient for the specific domain.

However, the test results do not outperform other models on every task. Clef-omni scored 98.2% on the BFCL test for matching cases, 92.7% on API-Bank, and 94.8% on BANKING77, but it fell below Clef and Clef-flash in some tests, such as Home Devices, When2Call, and PhishNChips. Therefore, the addition of multimodal capabilities does not mean overall superiority in every scenario.

Clef-flash Price Cut and Clef Speed Boost

Cloudflare reduced Clef-flash’s price from $0.09 to $0.038 per million input tokens, while Clef’s price remains $0.24, and Clef-omni launched at $0.15 per million input tokens. At the same time, the context window in the hosted version of Clef-flash was reduced from 64,000 to 24,000 tokens. The company says that only 0.24% of requests exceed the new limit, while the Hugging Face weights remain trained to support a window of up to 256,000 tokens when self-hosted.

Cloudflare also improved the speed of hosted Clef on Workers AI. Median latency fell from 262 to 152 milliseconds for inputs of approximately 800 tokens, and from 616 to 305 milliseconds for inputs of approximately 3,400 tokens, representing a speedup of up to 2.0 times. The company attributes part of the improvement to the use of SGLang, with contributions set to appear in version 0.5.22.

Why Does This Announcement Matter?

The announcement presents a practical direction for decision models: processing different inputs and extracting structured options instead of producing long text. Teams within Cloudflare have used these models to detect spam messages in GitHub, inspect EmDash system extensions for phishing, monitor personal metadata, and detect malicious domains. However, these examples are internal, and the mixed test results and context-window limitations in Clef-flash remain factors that should be evaluated before adopting the model in sensitive operations.

News source
Cloudflare Blog
Open original source ↗
c
Author

certi.news Editorial Team

In the same category

You may also like

View all news