@Cloudflare (2026-10-01, #BirthdayWeek) launched its first in-house decision models Clef (27B) and Clef-flash (9B): hosted on Workers AI, Apache 2.0 weights, Jev System One–compatible; official blog cites ~2.5× / 13× median latency vs Jev and 2.2s vs 4.7s in a threat-intel Browser Run workflow.

Key Takeaways

  • ✓Primary: Cloudflare 2026-10-01 → Clef blog
  • ✓Median latency (43 runs): Clef 209.3 ms, Clef-flash 38.8 ms vs Jev 524.1 ms (~2.5× / 13×)
  • ✓Quality: BFCL 98.47/98.76; BANKING77 94.20; CLINC150+OOS 97.43; Clef family tops 7/10 decision benches
  • ✓Shape: Qwen backbone + non-AR schema scoring; 64K ctx; vision; @cf/cloudflare/clef / clef-flash; $0.24 / $0.09 per M input tokens
  • ✓Ship: Workers AI binding/REST; weights on HF; RL fine-tune via design-partner form
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points

@Cloudflare on 2026-10-01 (#BirthdayWeek) pointed to Introducing Clef: the Workers AI team’s first in-house decision models Clef and Clef-flash. Agent hot paths need cheap, fast, structured classify/route answers—not another autoregressive round. After Typesafe’s Jev popularized the pattern, Cloudflare ships a System One–compatible open alternative for its “agent cloud” thesis.

Architecture Highlights & Internals

A decision model takes state plus up to 64 typed questions (noul / choice / score) and returns per-option probabilities—no free-form text to parse. Per changelog: Clef 27B, Clef-flash 9B, 64K context, vision encoder (≤4 images). Backbone: frozen Qwen (Qwen3.8-27B / Qwen3.5-9B) + rank-256 LoRA and routing head; inference is prefill-only then parallel schema scoring (non-autoregressive). RL fine-tune path: AI Gateway datasets → Containers sandbox → Trainer → BYO Model on Workers AI.

Authoritative Benchmarks & Measured Scores

All figures are Cloudflare’s own (blog/changelog), not third-party replications: median latency Clef 209.3 ms, flash 38.8 ms, Jev 524.1 ms (~2.5× / 13×); p95 238.6 / 122.4 / 536.0 ms. Quality excerpts: BFCL 98.47 / 98.76 (vs Jev 95.75); BANKING77 macro-F1 94.20; CLINC150+OOS 97.43; Home appliances flash 97.73. Threat intel + Browser Run: Clef 2.2 s vs gpt-oss-120b 4.7 s. On Typesafe workflow evals, Clef wins 3/4 vs Jev. Full tables: blog + HF card.

Developer Hands-on Guide

  1. Workers: env.AI.run("@cf/cloudflare/clef", { model: "clef", state, questions }); REST /ai/run/@cf/cloudflare/clef — see clef / clef-flash.
  2. Pricing (docs): $0.24 / $0.09 per M input tokens; enterprise default: no read/store/train on your traffic unless fine-tuning.
  3. Apache 2.0 weights: Cloudflare/clef, Cloudflare/clef-flash.
  4. Domain fine-tune: RL design-partner form; self-serve platform still maturing. Treat the tweet as a pointer—blog/docs are the SLA.