← Use cases

JevBench

JevBench is the shorthand for the independent benchmarks that have sprung up to rank Jev-class decision models — you hand the model some state and a bounded rubric, and it returns a typed answer with a probability for each option. Here's what the leaderboards measure, where Jev lands, and why to read the numbers carefully. These are third-party projects — not TypeSafe, and not us.

A decision model is easy to benchmark badly and hard to benchmark well: accuracy alone misses the point, because the whole value is a calibrated, typed answer returned cheaply and fast. The JevBench efforts try to capture that — scoring typed-decision accuracy, calibration, latency and cost across a fixed task set, and in some cases head-to-head games. Because several groups run their own, 'the JevBench score' depends on which board you mean.

The main leaderboards

LeaderboardWhat it measuresReported top of the table
Benchmark Heaven (JevBench v1.5.4)Typed-decision accuracy on a fixed rubric setCygnet ~73.7, Jev 1.13.0 ~72.1 (their numbers)
jevbench.devGames won head-to-head (StarCraft II, Minecraft)TypeSafe Jev 1.13 ~84% win rate (their numbers)
Various blog write-upsRe-runs of the public task setFigures around the mid-70s (unverified)

Treat every figure as reported-by-that-board, not gospel: the task sets differ, the runs are self-published, and none are audited. The useful takeaway isn't a single number — it's that Jev-class models now cluster closely on accuracy, so the real differentiators are calibration, latency, cost and whether the score is pinned between versions.

How Jev (TypeSafe) is positioned

On TypeSafe's own four-workflow benchmark Jev reaches roughly 67.8% agreement with reference answers — on par with GPT-5.6 Terra (67.9%) while running about 25x faster (0.4s vs 10.1s) — and posts a 0% structured-output error rate where general LLMs land anywhere from 0.58% to 45.5%.

On the independent boards Jev tends to sit at or near the top of the accuracy cluster while leading on speed, and its distinguishing claim is calibration: trained with RLCD so an 0.8 means roughly 80% across many calls, and pinned so the number doesn't drift between deploys. A benchmark can rank accuracy; what matters in production is whether you can trust a threshold like 'escalate if p > 0.8' — which is calibration, not a leaderboard rank.

Benchmark it yourself

The only numbers that matter are yours, on your data. Run your own decisions through the browser playground to feel the latency and calibration, then point your labelled set at the hosted endpoint and measure accuracy and calibration on the task you actually care about.

FAQ

What is JevBench?

JevBench is the name used by several independent benchmarks that rank Jev-class decision models — models that take state plus a bounded rubric and return a typed answer with a probability per option. They score accuracy, and in some cases calibration, latency, cost or head-to-head games. They're third-party, not built or endorsed by TypeSafe or by us.

What's the JevBench score for Jev?

It depends on the board. Benchmark Heaven's JevBench v1.5.4 reports Jev 1.13.0 around 72.1 (with Cygnet just ahead at ~73.7); jevbench.dev reports TypeSafe Jev 1.13 at about an 84% win rate in its games. Treat each as reported-by-that-board — the task sets differ and the runs are self-published.

Is JevBench official?

No. The leaderboards are independent community and third-party projects. Jev is one of the models they measure, but TypeSafe doesn't run them, and neither do we — this page just explains the landscape and links the sources.

How should I read the numbers?

Carefully. Jev-class models cluster closely on accuracy, so a one- or two-point gap between leaderboard rows rarely decides anything. What separates them in production is calibration (does an 0.8 mean 80%?), latency, cost, and whether the score is pinned between versions — so benchmark on your own data before choosing.

How do I benchmark Jev myself?

Try decisions free in the browser playground to feel the latency and calibration, then run your own labelled set against the hosted endpoint and measure accuracy and calibration on your actual task. Those numbers matter more than any public leaderboard.

See also: Jev's own benchmark numbers · Jev alternatives · Awesome Jev · Playground

The only benchmark that matters is yours

Run your own decisions free in the browser, then grab a jv_live_ key and measure Jev on your data.

▶ Try Jev freeGet an API key →
JevBench — the benchmarks for Jev-class decision models · Jev by TypeSafe AI