Clef & Clef-flash vs Jev
Clef and Clef-flash are Cloudflare's open-weight decision models, launched on Workers AI on 1 October 2026. They do the same thing Jev does — take application state plus a schema of typed questions and return a calibrated probability for each option in a single pass, no prose. That makes them the first serious open-weight challengers in the category Jev defined. Here's the honest comparison: what they are, how Cloudflare's own numbers stack up, and when open-weight beats hosted.
If you're comparing Clef-flash to Jev, you're choosing between two routes to the same kind of answer: a hosted, managed decision model (Jev) versus open weights you run yourself or call through Cloudflare's edge (Clef / Clef-flash). Both return typed, calibrated decisions — the trade-off is ops, calibration and where the model runs, not the shape of the call.
What Clef and Clef-flash are
Clef is Cloudflare's decision model, with Clef-flash as the smaller, latency-focused sibling. Cloudflare describes them as its first in-house open-weight decision models: instead of generating free-form text, each returns strictly typed outputs — a probability for every allowed option of every question you send — in one forward pass. Both are multimodal (text, JSON, and up to four images per request) and both ship as open weights on Hugging Face under Apache-2.0, so you can run them locally or call them on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. Per Cloudflare's post they freeze Qwen as the backbone — Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash — and add a vision encoder.
Specs, side by side
Cloudflare publishes the Workers AI pricing and context figures; the Jev column uses TypeSafe's own published figures. Treat every vendor's self-reported numbers as exactly that.
| Clef | Clef-flash | Jev | |
|---|---|---|---|
| Vendor | Cloudflare | Cloudflare | TypeSafe AI |
| Backbone | Qwen3.8-27B (frozen) | Qwen3.5-9B (frozen) | Purpose-built System One |
| Weights | Open · Apache-2.0 | Open · Apache-2.0 | Hosted (not open) |
| How you run it | Workers AI or self-host | Workers AI or self-host | Hosted API, self-serve key |
| Output | Typed probabilities | Typed probabilities | choice / score / noul + confidence |
| Multimodal | Text + up to 4 images | Text + up to 4 images | Text / JSON state |
| Context (vendor) | 64k | 64k | ~64k req / ~32k per question |
| Input price (vendor) | $0.24 / M | $0.09 / M | $0.25–$0.42 / M (output free) |
| Calibration | RL fine-tuned | RL fine-tuned | RLCD-calibrated, pinned versions |
Cloudflare's benchmark claims — read them honestly
Cloudflare's launch post benchmarks Clef against Jev and reports wins on several tasks: on BANKING77 (macro-F1) it lists Clef 94.20 vs Jev 79.74, on CLINC150+OOS 97.43 vs 89.27, and on BFCL (case exact) 98.47 vs 95.75. It also says Clef beat Jev in three of four areas on TypeSafe's own workflow evals, and puts Clef's median latency at 209.3ms against Jev's 524.1ms. These are Cloudflare's own figures, not an independent rerun.
The same post also shows Jev ahead on some tasks — When2Call (80.97 vs 72.37), BRIGHT (47.52 vs 45.91), and agent-trace observability (71.6 vs 68.5) — so even by Cloudflare's numbers it's task-dependent, not a clean sweep. TypeSafe separately reports Jev at roughly 70–500ms per call, below Cloudflare's 524.1ms figure, which is a useful reminder that latency depends heavily on who's measuring and how. If the decision is load-bearing, benchmark both on your own data before trusting any vendor's table.
When to pick which
- Reach for Clef or Clef-flash when you want open weights — to self-host for data residency or air-gapped runs, to fine-tune on your own data with Cloudflare's RL service, or because your stack already lives on Workers and an edge call is the natural fit.
- Reach for Jev when you want a hosted, managed decision model with calibrated (RLCD) confidence and pinned versions, no GPUs or inference ops to run, and a drop-in API you can point code at today — plus a free browser playground to feel the behaviour before you write a line.
- The call shape is nearly identical — state plus typed questions in, probabilities out — so prototyping on one and keeping the other as a fallback is cheap. The real decision is open-weight-and-self-managed vs hosted-and-calibrated.
Try the hosted side in one call
Whichever way you lean, the fastest way to feel the category is to make a real decision. Jev's browser playground does it with no signup, and a jv_live_ key points your own code at the hosted endpoint:
// A typed decision from the hosted model — same shape an open-weight
// decision model expects: state + typed questions in, probabilities out.
const res = await fetch("https://jevtypesafeai.com/api/v1/decide", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.JEV_KEY}`, // jv_live_...
},
body: JSON.stringify({
state: { ticket: ticket.body },
questions: {
queue: { type: "choice", instructions: "Route this ticket.",
criteria: { billing: "payments", bug: "something broken", sales: "pricing" } },
urgent: { type: "noul", instructions: "Is it urgent?" },
},
}),
});
const { answers } = await res.json();
route(answers.queue.choice, answers.urgent.noul > 0.8); // calibrated, sub-cent- Cloudflare blog: Introducing Clef — Cloudflare's launch post for Clef / Clef-flash, with its own benchmark table (vendor-reported)
- Workers AI model: clef-flash — Cloudflare's model page — 9B decision model, 64k context, $0.09/M input
FAQ
What is Clef-flash?
Clef-flash is the smaller of Cloudflare's two open-weight decision models (the larger is Clef), launched on Workers AI on 1 October 2026. Built on a frozen Qwen3.5-9B backbone with a vision encoder, it returns a calibrated probability for each option of each typed question in a single pass instead of generating text — the same job Jev does, as open weights.
Is Clef the same kind of model as Jev?
Yes — both are decision models: you send state plus typed questions and get back probabilities over your options, not prose. The difference is that Clef and Clef-flash are open-weight (Apache-2.0) models you self-host or call on Cloudflare Workers AI, while Jev is a hosted, RLCD-calibrated model you reach through a managed API.
Is Clef better than Jev?
Cloudflare's own launch benchmarks show Clef ahead on several classification tasks (e.g. BANKING77, CLINC150+OOS) and with lower median latency, but Jev leading on others (When2Call, BRIGHT, agent-trace observability). Those are vendor-reported numbers, not an independent rerun, so 'better' is task-dependent — benchmark both on your own data before deciding.
How much does Clef-flash cost vs Jev?
On Workers AI, Cloudflare lists Clef-flash at $0.09 per million input tokens and Clef at $0.24, output not billed. Jev bills per input token with output free as well; TypeSafe prices it by usage (roughly $0.25–$0.42 per million input depending on pack size) — see the pricing page for the current numbers. Open weights also carry your own inference/ops cost if you self-host.
Can I switch between Clef and Jev?
The call shape is nearly identical — application state plus a schema of typed questions, calibrated probabilities out — so moving a prototype between an open-weight model like Clef and hosted Jev is mostly a transport change. Many teams prototype on one and keep the other as a fallback.
See also: What is a decision model · Jev alternatives · Perplexity Decider vs Jev · Jev on Cloudflare Workers · Playground
Hosted and calibrated, or open and self-run
Clef gives you open weights; Jev gives you a hosted, calibrated decision in one call. Try Jev free in the browser, then grab a jv_live_ key.