Jev on vLLM
If you're searching for 'vllm jev', the honest answer is that hosted Jev (TypeSafe) is a closed, calibrated model with its own endpoint — there are no weights to load into vLLM. But the thing you actually want — a typed decision server you run on your own GPU — exists: OpenJev is a community project that speaks the same /v1/systemone wire format as Jev and serves an open model through vLLM, so your existing Jev SDK works against it by changing one base URL.
OpenJev is an open-source, Apache-2.0 "System One" decision server, not built by TypeSafe or by us. It takes a state and typed questions (noul / choice / score) and returns a probability and confidence for each in tens of milliseconds — the same request and response shape as hosted Jev, so the TypeSafe SDKs work against it unchanged. Its default backend runs DiffusionGemma 26B-A4B through vLLM on an NVIDIA GPU (24GB+), or through MLX on Apple silicon; it can also serve smaller open encoder models (Laya, Verdict, CLM, JevK5) when you want something CPU-friendly.
Serve it on vLLM (NVIDIA GPU)
The Docker path brings up OpenJev with its vLLM backend and listens on port 8080:
git clone https://github.com/razorback16/openjev && cd openjev
docker compose up -d
# OpenJev now serves /v1/systemone on 127.0.0.1:8080, DiffusionGemma 26B-A4B via vLLM
# On Apple silicon instead: pip install -e '.[mlx]' && OPENJEV_BACKEND=mlx python -m openjevIt accepts Jev-compatible model names (jev-latest, jev-preview) alongside its own (openjev-latest, diffusiongemma-26b), so code written for hosted Jev points at the local server by changing only the base URL:
// Same request shape as hosted Jev — just a local base URL, no key needed
const res = await fetch("http://localhost:8080/v1/systemone", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
model: "jev-latest",
state: { ticket: text },
questions: {
is_billing: { type: "noul", instructions: "Is this a billing issue?" },
},
}),
});
const { answers } = await res.json();
if (answers.is_billing.noul) routeToBilling(); // same typed shape hosted Jev returnsWhat you give up self-hosting on vLLM
| OpenJev on vLLM | Hosted Jev | |
|---|---|---|
| Weights | Open (DiffusionGemma, Apache-2.0) | Closed, versioned & pinned |
| Runs on | Your NVIDIA GPU (24GB+) or MLX | Hosted API |
| Calibration | Not RLCD-calibrated — you validate | RLCD-calibrated |
| Setup | Clone, docker compose up, own the GPU | Grab a key |
| Cost | Free to run, you pay for the GPU | Sub-cent per decision, output free |
| Best for | Private, offline, air-gapped, bulk | Production thresholds, zero ops |
The clean split: serve OpenJev on vLLM when you need decisions to stay on your own hardware — private, offline, or air-gapped — and you're willing to own GPU ops and validation. Call hosted Jev when you want a calibrated, pinned decision you can trust a production threshold against with nothing to run. Both answer the same typed shape — choice, score, noul — so the interface your code depends on never changes.
FAQ
Can I run Jev on vLLM?
Not hosted Jev itself — its weights are closed, so there's nothing to load into vLLM. But OpenJev, a community project, serves a Jev-wire-compatible model (DiffusionGemma 26B-A4B) through vLLM on an NVIDIA GPU, and because it speaks the same /v1/systemone format your Jev SDK works against it by changing one base URL.
What hardware does OpenJev's vLLM backend need?
Its DiffusionGemma backend wants an NVIDIA GPU with 24GB+ of VRAM. On Apple silicon it runs through MLX instead, and it can also serve smaller open encoder models (Laya, Verdict, CLM, JevK5) when you want something lighter or CPU-friendly.
Is a self-hosted model as good as hosted Jev?
Not out of the box. OpenJev's open weights aren't RLCD-calibrated, so a confidence of 0.8 isn't guaranteed to mean 80% until you validate it on your own labelled data. Hosted Jev is calibrated and pinned, which is what you're paying for when you need to trust a threshold.
Is OpenJev official?
No — OpenJev is a community, open-source project (Apache-2.0), not built by TypeSafe AI or by us. It's compatible with the Jev wire API by design, which is why the SDKs work against it, but it's a separate implementation on open weights.
See also: Run a Jev-like model with Ollama · Jev alternatives · Run a decision model locally · Pricing
Want calibrated, with zero GPU ops?
Self-hosting on vLLM is great when the data has to stay on your box. For a calibrated decision with nothing to run, try hosted Jev free in the browser, then grab a jv_live_ key.