CLOUDFLARE has announced Clef, a new family of open‑source decision models, and Clef-flash, with an accompanying RL fine‑tuning platform. The models, hosted on Workers AI and released under Apache 2.0 on Hugging Face, are designed to produce deterministic, structured outputs that can be directly integrated into decision workflows.
Clef is positioned as the leading model in the Jev decision index among Cloudflare’s offerings, and both Clef and Clef-flash are fully API‑compatible with Jev, enabling easy substitution in existing pipelines. Cloudflare also highlights that the models can ingest images (vision encoder) in addition to text, and operate with a 64k context window, which is larger than Jev’s 32k.
Benchmarks reported by Cloudflare show Clef models delivering superior latency in 43 evaluated tasks, with Clef‑flash delivering notably faster results than the larger Clef model and other competitors. In practical terms, Clef achieved a median latency of 209.3 ms (Clef) and 38.8 ms (Clef‑flash) in their tests, compared with 524.1 ms for Jev. The company notes that hosting Clef on Cloudflare’s edge GPUs via Workers AI reduces network latency, making Clef viable for hot‑path agential decisions.
The article also details a two‑stage inference approach using Qwen as the base model, plus reinforcement learning for calibrated decisions (RLCD), to improve accuracy and calibration. Beyond the public models, Cloudflare offers a hands‑on fine‑tuning service via its forward‑deployed engineering team, with plans for a self‑serve platform to train and redeploy tuned models on Workers AI.