Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instead of Text


Cloudflare has released Clef and Clef-flash, the first models trained by its Workers AI team. They are decision models, not chatbots. Each reads an input state and a schema of typed questions. It returns a probability for every allowed answer, with no free-form text. Both are open-weight under Apache 2.0 and compatible with TypeSafe AI’s Jev API.

Is it deployable? Yes, Both models run today on Workers AI, and the weights are on Hugging Face for self-hosting.

What a Decision Model Does

An LLM generates tokens one at a time, and its output still needs parsing. A decision model only answers a fixed set of questions about an input. Clef supports 3 question types:

  • noul: yes/no, returns the probability of yes.
  • choice: picks 1 named option, with per-option probabilities and a confidence value.
  • score: rates against an ordered rubric, returning a probability-weighted score.

On Workers AI, 1 request carries up to 64 questions and up to 4 images.

TypeSafe AI launched Jev, its first ‘System One’ model, on September 15, 2026. Open alternatives like Kev-9B and Laya followed. Clef uses the same System One API. Switching from Jev means changing the endpoint and model name.

How Clef Works

Clef is post-trained from Qwen3.8-27B, and Clef-flash from Qwen3.5-9B. Both keep the backbone’s vision encoder.

Inference has 2 stages. The backbone first runs a single prefill-only pass over the state and questions. A small transformer, the joint schema head, then reads the final hidden states. It routes evidence to each question, lets fields cross-attend, and scores all options jointly. A per-question softmax turns logits into probabilities.

Training froze both backbones and jointly optimized the routing head with rank-256 low-rank adapters. The loss pairs label-smoothed cross-entropy with a Brier loss for calibration. A secondary objective, Reinforcement Learning for Calibrated Decisions (RLCD), gives partial credit to adjacent ordinal choices.

Interactive Explainer



Source link

  • Related Posts

    A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End

    In this tutorial, we implement Kauldron, the JAX training library from Google Research that describes itself as optimized for research velocity and modularity, and we take those two words literally…

    Cohere Releases Embed 5: How It Compares to Voyage 4 Large, Gemini Embedding 2, and OpenAI

    Cohere has released Embed 5, a new embedding model family. It targets enterprise search, RAG, and agentic retrieval. The model family ships in 2 tiers. Embed 5 Pro targets maximum…

    Leave a Reply

    Your email address will not be published. Required fields are marked *