CLOUDFLARE has expanded its Clef family with Clef-omni, a multimodal decision model that now accepts audio, video, image, and text inputs in a single API call. Clef-omni processes wav or mp3 audio and mp4/webm video alongside images and textual data, removing the need for separate transcription or multi-model pipelines. The model is built on a Qwen3-Omni-30B-A3B-Instruct MoE foundation, with the primary comprehension backbone retained while discarding text-to-speech output components.
In practice, Clef-omni scores all modalities in one pass, pulling candidate values from internal embeddings and using a two-stage attention routing to compute confidence scores across the full input context. Output tokens are not produced because Clef models are not LLMs; instead, the system returns fast, schema-constrained decisions.
Cloudflare also reports faster performance and a pricing refresh across Clef models. Clef-omni delivers fast results: text-only decisions around 130 ms median, images about 150 ms, audio in a few hundred milliseconds, and a full 21-second video around 1.5 seconds in a single API call. The Clef-flash pricing has been reduced to $0.038 per million input tokens (down from $0.09), while Clef remains at $0.24 per million input tokens and Clef-omni at $0.15 per million input tokens.
To support cheaper usage, the hosted Clef-flash context window has been trimmed to 24k tokens (the Hugging Face weights themselves retain a 256k capability for self-hosts). Cloudflare notes that only a small fraction of requests would exceed 24k, and suggests users needing larger contexts consider Clef instead of Clef-flash. The speeds for Clef have also improved thanks to serving-layer optimisations and the adoption of SGLang, with forthcoming releases in SGLang 0.5.22.