Clef vs Jev
Clef and Jev are the two ends of the decision-model market in late 2026: Cloudflare's open-weight Clef (27B) and Clef-flash (9B) that you self-host or run on Workers AI, versus TypeSafe's hosted, RLCD-calibrated Jev that you reach through a managed API. Both take state plus typed questions and return calibrated probabilities in one pass — so this is an ops-and-calibration decision, not a 'can it decide' one. Here's the honest head-to-head.
If you're choosing between Clef and Jev, you've already decided you want a decision model rather than an LLM you prompt and parse. What's left is the trade-off: open weights you manage (Clef) versus a hosted, calibrated endpoint you call (Jev). The decide call looks almost the same either way.
Specs, side by side
The Clef columns use Cloudflare's published Workers AI figures; the Jev column uses TypeSafe's. Every number here is vendor-reported, not an independent rerun.
| Clef | Clef-flash | Jev | |
|---|---|---|---|
| Vendor | Cloudflare | Cloudflare | TypeSafe AI |
| Backbone | Qwen3.8-27B (frozen) | Qwen3.5-9B (frozen) | Purpose-built System One |
| Weights | Open · Apache-2.0 | Open · Apache-2.0 | Hosted (not open) |
| How you run it | Workers AI or self-host | Workers AI or self-host | Hosted API, self-serve key |
| Output | Typed probabilities | Typed probabilities | choice / score / noul + confidence |
| Multimodal | Text + up to 4 images | Text + up to 4 images | Text / JSON state |
| Input price (vendor) | $0.24 / M | $0.09 / M | $0.25–$0.42 / M (output free) |
| Calibration | RL fine-tuned | RL fine-tuned | RLCD-calibrated, pinned versions |
The benchmark claims — read them honestly
Cloudflare's launch post benchmarks Clef against Jev and reports wins on several classification tasks: BANKING77 macro-F1 94.20 vs 79.74, CLINC150+OOS 97.43 vs 89.27, BFCL case-exact 98.47 vs 95.75, and a median latency of 209.3ms against Jev's 524.1ms. It also says Clef beat Jev in three of four areas on TypeSafe's own workflow evals.
But the same post shows Jev ahead on others — When2Call 80.97 vs 72.37, BRIGHT 47.52 vs 45.91, agent-trace observability 71.6 vs 68.5 — so even by Cloudflare's numbers it's task-dependent, not a sweep. TypeSafe separately reports Jev at roughly 70–500ms per call, below Cloudflare's 524.1ms figure, which is a good reminder that latency depends on who's measuring. The only benchmark that settles your case is your decision, on your data.
When to pick which
- Pick Clef or Clef-flash when you want open weights — to self-host for data residency or air-gapped runs, to fine-tune on your own decisions with Cloudflare's RL service, or because your stack already lives on Workers and an edge call is the natural fit. Clef-flash if latency rules; Clef if accuracy does.
- Pick Jev when you want a hosted, managed model with RLCD-calibrated confidence and pinned versions, no GPUs or inference ops to run, image-free text/JSON decisions, and a drop-in API plus a free browser playground to feel the behaviour before you integrate.
- Not sure? Prototype on one and keep the other as a fallback — the call shape barely changes between them.
Migrating between them
Because both speak the same decide shape — state plus a schema of typed questions in, calibrated probabilities out — moving between Clef and Jev is mostly a transport change. Going Clef → Jev, you drop the model hosting and point at the hosted endpoint below, gaining RLCD calibration and pinned versions; going Jev → Clef, you take on inference ops in exchange for open weights and edge locality. Re-tune any confidence thresholds after a switch: calibration differs between a from-scratch RLCD model and an RL-fine-tuned open backbone, so an 0.8 cut-off may not mean the same thing on both.
Make one real decision on Jev
The fastest way to judge is to run your own decision through it. Jev's playground does it in the browser with no signup, and a jv_live_ key points your code at the hosted endpoint:
// Hosted, calibrated decision — the same state+questions shape Clef expects.
const res = await fetch("https://jevtypesafeai.com/api/v1/decide", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.JEV_KEY}`, // jv_live_...
},
body: JSON.stringify({
state: { ticket: ticket.body },
questions: {
queue: { type: "choice", instructions: "Route this ticket.",
criteria: { billing: "payments", bug: "something broken", sales: "pricing" } },
urgent: { type: "noul", instructions: "Is it urgent?" },
},
}),
});
const { answers } = await res.json();
route(answers.queue.choice, answers.urgent.noul > 0.8); // calibrated, sub-cent- Cloudflare blog: Introducing Clef — Cloudflare's launch post with its Clef-vs-Jev benchmark table (vendor-reported)
- Workers AI model: clef-flash — Cloudflare's model page — 9B decision model, 64k context, $0.09/M input
FAQ
What's the difference between Clef and Jev?
Clef is Cloudflare's open-weight decision model (Clef 27B and Clef-flash 9B) that you self-host or run on Workers AI; Jev is TypeSafe's hosted, RLCD-calibrated decision model with pinned versions. Both take state plus typed questions and return calibrated probabilities in one pass — the difference is open-weight-and-self-managed vs hosted-and-calibrated, not the shape of the call.
Is Clef better than Jev?
Cloudflare's own launch benchmarks show Clef ahead on several classification tasks (BANKING77, CLINC150+OOS, BFCL) and with lower median latency, but Jev leading on others (When2Call, BRIGHT, agent-trace observability). Those are vendor-reported numbers, not an independent rerun, so 'better' is task-dependent — test both on your own data.
Is Clef cheaper than Jev?
On Workers AI, Cloudflare lists Clef-flash at $0.09 per million input tokens and Clef at $0.24, output not billed. Jev also bills per input token with output free; TypeSafe prices it by usage, roughly $0.25–$0.42 per million input depending on pack size. If you self-host Clef's open weights, add your own inference and ops cost.
Can I migrate from Clef to Jev (or back)?
Yes, with little rework — the decide shape is nearly identical, so it's mostly a transport change. Re-tune your confidence thresholds after switching, since calibration differs between Jev's from-scratch RLCD training and Clef's RL-fine-tuned open backbone.
Does Jev support images like Clef?
Not today — Jev reads text and JSON state, while Clef and Clef-flash accept up to four images per request. If your decision depends on image input, that's a point in Clef's favour; for text and structured state, Jev's hosted, calibrated route is the simpler one.
See also: What is Clef · Clef-flash specifically · Jev alternatives · What is a decision model · Playground
Open-weight, or hosted and calibrated
Clef gives you open weights; Jev gives you a hosted, calibrated decision in one call. Try Jev free in the browser, then grab a jv_live_ key.