Cloudflare has released Clef and Clef-flash, the first models trained directly by its Workers AI engineering team. Positioned as specialized “decision models” rather than conversational chatbots, the family is engineered to serve as high-speed routing engines, guardrails, and classifiers for autonomous AI agent pipelines.
Instead of auto-regressively generating text tokens that downstream systems must parse, extract, and validate, Clef processes an input state along with a structured schema of typed questions, returning exact, calibrated probabilities for predefined answer options in a single forward pass. Both models are distributed under the Apache 2.0 license on Hugging Face, run natively across Cloudflare’s global Workers AI GPU footprint, and provide drop-in compatibility with the API standard popularized by TypeSafe AI’s proprietary Jev decision model.
Key Takeaways
- Decision Models Over Chatbots: Clef generates no free-form text, markdown, or chain-of-thought tokens. It ingests an input state (text, JSON, images) and returns structured probabilities for up to 64 predefined questions per request.
- Low Latency on the Hot Path: Clef-flash achieves a median latency of 38.8 ms (p95 of 122.4 ms), while the larger Clef model logs a median latency of 209.3 ms, outpacing TypeSafe AI’s Jev (median 524.1 ms).
- Two Open-Weight Sizes:
- Clef (27B): Built on a frozen Qwen 3.8-27B backbone with a 64K-token context window; optimized for high-precision, multi-variable decisions.
- Clef-flash (9B): Built on a frozen Qwen 3.5-9B backbone; engineered for throughput-heavy, latency-sensitive gatekeeping.
- Multimodal Vision Ingestion: Unlike text-only decision alternatives, Clef supports image inputs alongside textual state data, enabling visual content moderation and UI inspection.
- Drop-In Compatibility with Jev: Implements the System One API specification, allowing teams to swap endpoints from Jev to Clef by changing the model identifier.
- Reinforcement Learning Fine-Tuning Service: Cloudflare launched an accompanying enterprise fine-tuning program using Reinforcement Learning for Calibrated Decisions (RLCD) to train custom decision heads on private organizational telemetry.
1. What Is a Decision Model, and Why Does It Matter for AI Agents?
In autonomous software architectures, large language models (LLMs) are frequently used not to write creative prose, but to make operational decisions: Should this support ticket escalate?, Which database tool should the agent run?, or Does this user query violate security policies?
Using general-purpose autoregressive models for these decisions introduces structural drawbacks:
THE HOT-PATH EXECUTION DILEMMA
│
┌──────────────────────────────┴──────────────────────────────┐
▼ ▼
TRADITIONAL LLM CALL (e.g., Llama / GPT) DECISION MODEL (Cloudflare Clef)
• Generates text token-by-token (slow) • Single prefill-only forward pass (fast)
• High latency (500ms to 3,000ms+) • Edge-native latency (39ms to 209ms)
• Output formatting can break or hallucinate • Output is strictly typed JSON probabilities
• Requires regex, JSON schemas, or retries • Direct if/else downstream code execution
Clef operates like a musical clef—setting the structural key for everything that follows. Downstream agent code can evaluate the returned confidence scores directly with standard logic operators:
TypeScript
// Example: Direct downstream execution based on Clef probability
if (decision.questions.urgent.probability > 0.85) {
await pagerduty.triggerIncident();
} else {
await helpdesk.routeToTeam(decision.questions.team.selected);
}
2. Technical Architecture: Two-Stage Inference and Joint Schema Heads
Clef is post-trained on top of Alibaba’s open Qwen architecture, preserving the backbone’s vision encoder while freezing base model weights:
+-----------------------------------------------------------------------------------+
| CLOUDFLARE CLEF: MODEL FAMILY SPECIFICATIONS |
+-----------------------------------------------------------------------------------+
| Technical Metric | Clef-flash | Clef (Full) |
+--------------------------------+--------------------------+-----------------------+
| **Parameter Count** | 9 Billion (Qwen 3.5-9B) | 27 Billion (Qwen 3.8) |
| **Median Latency** | **38.8 ms** | **209.3 ms** |
| **p95 Latency** | 122.4 ms | 238.6 ms |
| **Context Window** | 64,000 tokens | 64,000 tokens |
| **Modalities Supported** | Text, JSON, Images | Text, JSON, Images |
| **License** | Apache 2.0 (Open Weights)| Apache 2.0 |
| **API Compatibility** | TypeSafe AI Jev | TypeSafe AI Jev |
+--------------------------------+--------------------------+-----------------------+
TWO-STAGE INFERENCE PIPELINE
│
▼
1. STATE + SCHEMA INPUT (UP TO 64K TOKENS)
(Text, raw logs, JSON objects, up to 4 images)
│
▼
2. FROZEN QWEN BACKBONE PREFILL PASS
(Computes rich hidden-state representations)
│
▼
3. JOINT SCHEMA HEAD (RANK-256 LORA ADAPTER)
(Cross-attends question fields and scores options)
│
▼
4. PER-QUESTION SOFTMAX PROBABILITIES
(Strictly typed values, 0.0 to 1.0)
Supported Question Primitives
Clef accepts up to 64 simultaneous questions within a single request, categorizing them across three distinct typed question formats:
noul(Boolean / Yes-No): Returns a single probability float indicating the likelihood that the answer is “yes.”choice(Categorical): Evaluates mutually exclusive options (e.g., routing categories: Billing, Technical, Sales), outputting individual probabilities that sum to 1.0 alongside the top-selected label.score(Ordinal Scale): Evaluates an input against an ordered rubric, returning a probability-weighted scalar score.
3. Benchmark Comparisons: Clef vs. TypeSafe AI Jev
Cloudflare evaluated Clef against TypeSafe AI’s proprietary Jev across the 10-benchmark shortlist from the Decision Index 0.2.1 suite:
+-----------------------------------------------------------------------------------+
| DECISION BENCHMARKS: CLEF VS. JEV COMPARISON |
+-----------------------------------------------------------------------------------+
| Benchmark Suite / Domain | Clef (27B) | Clef-flash (9B) | Jev (Baseline) |
+--------------------------------+-------------------+------------------+-------------------+
| **Median Inference Latency** | **209.3 ms** | **38.8 ms** | 524.1 ms |
| **BANKING77 (Macro-F1)** | **94.20** | 90.93 | 79.74 |
| **CLINC150+OOS (Macro-F1)** | **97.43** | 66.77 | 89.27 |
| **BFCL Function Calling** | 98.47 | **98.76** | 95.75 |
| **Home Appliances (Exact)** | 82.95 | **97.73** | 52.27 |
+--------------------------------+-------------------+------------------+-------------------+
| **TypeSafe Real-World Evals** | | | |
| - Invoice Processing | **64.7** | — | 61.8 |
| - Customer Service Routing | **76.3** | — | 76.0 |
| - Security Incident Triage | **62.9** | — | 61.7 |
| - Agent Trace Observability | 68.5 | — | **71.6** |
+--------------------------------+-------------------+------------------+-------------------+
While Jev maintains a lead on deep, knowledge-retrieval academic tests (such as GPQA Diamond and MMLU-Pro), Clef demonstrates consistent advantages across high-frequency intent classification, function routing, and enterprise ticket handling. In an internal threat-intelligence pilot conducted by Cloudflare, pairing Clef with Browser Rendering fetched, parsed, and classified suspicious domain categories in 2.2 seconds, compared to 4.7 seconds using an open-weights 120B parameter conversational model.
4. Key Deployment Scenarios for Engineering Teams
AGENTIC USE CASES FOR CLEF
│
┌─────────────────────────────────┼─────────────────────────────────┐
▼ ▼ ▼
AGENT TOOL-CALL GUARDRAILS REAL-TIME TICKET ROUTING THREAT & CONTENT MODERATION
Verifies permissions & safety Directs incoming support tickets Scores user uploads & URLs
in <40ms before an agent invokes to departments without human against compliance rubrics
destructive production APIs. intervention or text parsing. with multimodal image support.
- Pre-Execution Tool Guardrails: Before an autonomous coding or administrative agent executes a sensitive shell script or database deletion, a sub-40ms call to Clef-flash evaluates: “Is this command safe and within operational scope?” If the probability drops below an acceptable safety threshold, the pipeline halts for human authorization.
- Deterministic Triage: High-volume customer service desks can classify unstructured user inquiries into specific queue categories, assign priority levels, and check for account churn risk in one round-trip.
- Multimodal Trust and Safety: Because Clef includes a vision encoder, teams can pass both user text and image attachments (up to four images per request) to flag policy violations, inappropriate imagery, or fraudulent document scans.
Availability and Deployment Options
- Workers AI: Both models are live across Cloudflare’s global edge network via the Workers AI binding (
@cf/cloudflare/clefand@cf/cloudflare/clef-flash) and accessible through REST endpoints at/ai/run. - Self-Hosting: Complete model weights and configuration files are hosted on Hugging Face under the Apache 2.0 license for deployment in private VPCs and on-premises clusters.
- RLCD Fine-Tuning Service: Cloudflare is accepting enterprise design partners for its managed reinforcement learning fine-tuning pipeline, leveraging acquired Replicate infrastructure to let organizations train custom decision heads on proprietary data.
Frequently Asked Questions (FAQs)
What is the difference between a decision model and a regular LLM?
A regular LLM generates text one token at a time, which can be slow and unpredictable in format. A decision model takes an input and a fixed schema of questions, running a single forward pass to return structured, calibrated probabilities for predefined options without producing conversational text.
How fast are Cloudflare’s Clef models?
Clef-flash achieves a median latency of 38.8 milliseconds, making it roughly 13 times faster than TypeSafe AI’s Jev. The larger 27B Clef model operates at a median latency of 209.3 milliseconds.
Are Clef and Clef-flash open-source?
Yes. Both models are released as open-weights under the permissive Apache 2.0 license on Hugging Face, allowing teams to run them locally or on private cloud infrastructure.
Can Clef process images?
Yes. Both Clef (27B) and Clef-flash (9B) retain the native vision encoder from their underlying Qwen backbones, enabling them to evaluate up to four images alongside text state per request.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



