Cloudflare has recently officially released the new open-weight decision model Clef-omni. This model represents a major upgrade based on the existing Clef series, supporting text and images natively, and for the first time, adding direct processing capabilities for audio and full video inputs. Currently, its model weights are fully open-sourced on the Hugging Face platform.

In terms of technical implementation and functional features, Clef-omni completely changes the traditional processing pipeline. Previously, when handling video inputs, Clef needed to split videos into continuous static images along the timeline before processing; while the new Clef-omni directly supports mainstream audio and video formats such as wav, mp3, mp4, and webm. This means developers no longer need to separately build complex speech transcription and audio-video splitting pipelines. They can simply use a single API call to achieve unified processing of text, images, audio, and video.

image.png

From the perspective of performance and benchmark test results, this model is built based on the Qwen3-Omni-30B-A3B-Instruct foundation architecture, retaining its core understanding capabilities. As a specialized model focused on structured decision-making tasks, it does not output conventional lengthy texts. According to official test data, the median response time for pure text requests is approximately 130 milliseconds, and for image requests around 150 milliseconds; processing a 21-second video with sound takes only about 1.5 seconds to complete the scoring efficiently.

Along with the release of the new model, Cloudflare has also made price adjustments and optimizations for other product lines. The input price of Clef-flash has been significantly reduced by about 58%, from the original 0.09 USD per million input tokens to 0.038 USD, and the context window for the hosted version has been adjusted from 64k to 24k simultaneously. Overall, with the open-source release of Clef-omni and the completion of multi-modal processing capabilities, it is providing a more efficient and cost-effective underlying support for global developers to build complex intelligent workflows and automated decision systems.