← Use cases

CLM-8B vs Jev

CLM-8B vs Jev is the newest 'is the open model faster than Jev?' question, and this time the headline is real: in its own tests CLM-8B runs up to 9x faster than Jev. But speed is one axis. CLM-8B is Stanford and Nvidia's open-weight 8B contrastive model that you run yourself; Jev (TypeSafe) is the hosted, closed-weight System One model trained and calibrated for decisions. Same job — pick the right typed answer — different trade-offs on accuracy, calibration and ops.

If you're searching "clm-8b vs jev" or "clm 8b", you've seen the framing going around after the Stanford/Nvidia release: an open 8B model that selects an agent's next action up to 9x faster than Jev. It's a genuinely interesting result and worth taking seriously — but the benchmark numbers cut both ways, so it's worth being precise about what each model is and where each one wins.

What CLM-8B is

CLM-8B is an open-weight, 8-billion-parameter Contrastive Language Model from researchers at Stanford and Nvidia. Instead of generating an answer token by token, it builds a vector representation of the current state and of each available action, then picks the action whose vector best matches the state — closer to a retrieval/ranking step than to text generation. It also caches reusable agent actions, which is a big part of why its decode is so fast. The weights are open, so you self-host it.

Why it's positioned against Jev

Because it targets exactly Jev's job: the choice step inside an agent loop. A typed decision — pick one of these actions, one of these labels — doesn't need left-to-right generation, it needs one structured pick. CLM-8B's contrastive match does that in a single fast pass locally, with open weights and no per-call bill. That's the same primitive Jev's choice covers, which is why the head-to-head appeared — the open, local cousin of a hosted decision model, in the same lineage as running Laya or NanoJev yourself.

CLM-8B vs Jev at a glance

Jev (TypeSafe)CLM-8B (Stanford/Nvidia)
What it isHosted System One decision modelOpen-weight 8B contrastive model
How it decidesNon-autoregressive typed outputVector match: state vs cached actions
Where it runsHosted APILocal / self-hosted (your GPU)
WeightsClosed, versioned & pinnedOpen-weight
Outputchoice · score · noul, typedBest-matching action (choice-shaped)
CalibrationRLCD-trained, calibrated probabilitiesNot calibrated for general use
Speed~70–500ms incl. networkUp to ~9x faster than Jev locally (its tests)
Accuracy (BFCL v4 tool-calling)99.2%95.2%
WikiRacing (tasks completed)30 / 3026 / 30
SetupMinutes — grab a keyHigher — weights, GPU, runtime
Cost~$0.001 / decision, output freeFree to run, you pay compute
Best forProduction, calibrated numbers, no opsLocal/offline, open weights, raw latency

About the '9x faster than Jev' number

It's real, and it's the honest headline of the CLM-8B release: in zero-shot tests across computer use, gaming and tool calling, CLM-8B ran up to 9x faster than Jev and matched its success rate on a couple of game tasks. That's a strong result for an open 8B model. But read the same benchmark table to the end: CLM-8B scored 95.2% to Jev's 99.2% on the BFCL v4 tool-calling benchmark, and finished 26 of 30 WikiRacing tasks to Jev's 30. So the trade is concrete — meaningfully faster, a few points less accurate, and a 4-point gap on tool calling is a lot when a wrong pick triggers the wrong action.

When CLM-8B wins

Reach for CLM-8B when raw local latency and open weights matter most and you can own quality yourself: high-volume decisions on your own hardware, offline or regulated environments, or research where you want to inspect and cache the action space directly. Its contrastive match is fast and cheap to run, and there's no per-call cost. The trade is the usual open-model one — you host the GPU and runtime, and because it isn't calibrated for general use, you validate accuracy and confidence on your own labelled data before wiring a threshold to it.

When Jev wins

Reach for hosted Jev when you want the most accurate, calibrated decision without owning a model. Jev leads on the accuracy benchmarks (99.2% vs 95.2% on BFCL v4, 30/30 vs 26/30 WikiRacing) and is trained with RLCD, so its probabilities are calibrated — an 0.8 means roughly 80% across many calls — and pinned, so the number doesn't drift between deploys. You also skip all inference ops. The moment 'escalate if p > 0.8' goes to production, that calibration is the difference: CLM-8B gives you a fast pick, Jev gives you a pick plus a number you can trust a threshold against.

Can you use both?

Yes, and it's a sensible split. Run CLM-8B locally for the highest-volume, latency-critical picks where you own the box and can validate quality, and call hosted Jev for the decisions where you want the calibrated, pinned contract and zero ops — the risk gates and routing where a 4-point accuracy gap or an uncalibrated number would cost you. Both answer the same shape — a typed choice — so the interface your code depends on can stay identical while you choose what answers it. The fastest way to feel the hosted side is the playground: a real calibrated decision in the browser, free, before you commit.

FAQ

Is CLM-8B really 9x faster than Jev?

In its own zero-shot tests, yes — up to 9x faster across computer use, gaming and tool calling, and it matched Jev's success rate on two game tasks. But the same benchmarks show it's less accurate: 95.2% vs Jev's 99.2% on BFCL v4 tool-calling, and 26 of 30 WikiRacing tasks to Jev's 30. Faster, a few points less accurate.

What is CLM-8B?

An open-weight, 8B Contrastive Language Model from Stanford and Nvidia. It picks an agent's action by matching a vector of the current state against vectors of the available (and cached) actions, rather than generating text — a fast, retrieval-like decision step you self-host.

Is CLM-8B calibrated like Jev?

No. Jev is trained with RLCD to return calibrated probabilities you can threshold. CLM-8B isn't calibrated for general use, so any confidence you read from its match scores should be validated on your own labelled data before you set thresholds against it.

Can CLM-8B replace hosted Jev in production?

It can for latency-critical, high-volume picks where you own the GPU and validate quality yourself. For decisions that need the highest accuracy and a calibrated, pinned probability with zero ops, hosted Jev still leads on the accuracy benchmarks and gives you a number you can trust a threshold against.

Which is cheaper?

CLM-8B is free to run once you own the hardware — you pay compute. Hosted Jev is about $0.001 per decision with output free and no GPU or ops. Which is cheaper depends on volume and whether you already run GPUs: at high steady volume self-hosting can win; at low or bursty volume the hosted call usually does.

See also: NanoJev vs Jev · Laya vs Jev · Running a decision model locally · Playground

Try the calibrated, hosted model free

CLM-8B is fast and open. For the most accurate, calibrated decisions with zero ops, run one free in the browser — no signup — then grab a jv_live_ key.

▶ Try Jev freeGet an API key →
CLM-8B vs Jev: the open 8B decision model vs the hosted System One · Jev by TypeSafe AI