The Clef family grows
Cloudflare has announced Clef-omni, following last week's release of its open-weight decision models Clef and Clef-flash. According to the company, the Clef idea came together in under a week: the model was trained over a weekend and shipped the same week.
What multimodality brings
Clef-omni accepts audio (wav or mp3) and video (mp4 or webm) input in addition to text and images. That means instead of building cascaded pipelines that transcribe speech to text or split audio and image channels from video, users can make decisions across modalities with a single model call.
Technical foundation
The model is built on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts (MoE) foundation. Cloudflare adopted the comprehension backbone while discarding the text-to-speech output components. Because Clef models are not large language models, no output tokens are generated, removing the overhead of transcribing or captioning incoming files. Training froze the Qwen3 backbone and trained low-rank adapters (LoRA).
Price and speed
Cloudflare said it cut the price of Clef-flash, making it cheaper than Jev, and made Clef faster. The model weights were released openly on HuggingFace.



