Jev on Ollama
If you're looking to 'ollama pull jev', the honest answer is that hosted Jev is a closed, calibrated model with its own endpoint — it isn't an Ollama model. But the thing you actually want — a typed decision model running locally — does exist: Ollaya is the 'Ollama for decision models', and it serves open Jev-like models behind a TypeSafe-compatible API.
Jev (TypeSafe) is hosted and closed-weight, so there's no GGUF to pull into Ollama. What the community built instead is Ollaya — a local runtime that pulls and serves open decision models (Laya, winnow, NLI, GLiClass and more) by name, the way Ollama serves LLMs. It speaks TypeSafe's /v1/systemone wire format, so code written for Jev points at it by changing one base-URL env var. These are community, open-weight models — not Jev, and not calibrated the way hosted Jev is — so validate any of them on your own data.
Run one locally
# Install Ollaya (the 'Ollama for decision models')
curl -fsSL https://ollaya.dev/install.sh | sh
# Pull + serve an open Jev-like model (Laya = ModernBERT-large, 421M, CPU-friendly)
ollaya run laya --preset triage
# GGUF models (winnow, jevk5) run on llama.cpp (CPU/CUDA/Metal); on Mac, laya runs on the Apple GPU via MLXOllaya exposes a TypeSafe-compatible endpoint locally, so the same decide call you'd send Jev works against it — you just point the base URL at the local daemon:
// Same request shape as hosted Jev — just a local base URL, no key
const res = await fetch("http://localhost:11435/v1/systemone", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
state: { ticket: text },
questions: {
queue: { type: "choice", instructions: "Which queue should this go to?",
criteria: { billing: "payments", technical: "bugs", abuse: "spam/fraud", sales: "pricing" } },
},
}),
});
const { answers } = await res.json();
route(answers.queue.choice); // one of your criteria keys — same shape Jev returnsWhat you give up running locally
- Calibration — open models aren't RLCD-calibrated, so a 0.8 doesn't reliably mean 80%; you own the validation
- Accuracy — the small open models trade points of accuracy for size and speed; check them on your labelled data
- Pinning & ops — you run the daemon, pick models, and keep them updated; nothing is versioned for you
- No confidence you can trust a production threshold against out of the box
When to run local vs call hosted Jev
| Ollaya (local) | Hosted Jev | |
|---|---|---|
| Weights | Open (Laya, winnow, …) | Closed, versioned & pinned |
| Runs on | Your CPU/GPU (llama.cpp / MLX) | Hosted API |
| Calibration | Not calibrated — you validate | RLCD-calibrated |
| Setup | Install daemon, pull a model | Grab a key |
| Cost | Free to run, you pay compute | Sub-cent per decision, output free |
| Best for | Offline, private, tinkering | Production volume, trusted thresholds, zero ops |
The clean split: run Ollaya locally for offline, private, or high-volume bulk decisions where you own the box, and call hosted Jev when you want a calibrated, pinned decision you can trust a threshold against with no ops. Both answer the same typed shape — choice, score, noul — so the interface your code depends on stays identical. See the local-model guide for the full self-host path.
FAQ
Can I run Jev with Ollama?
Not hosted Jev itself — it's a closed, calibrated model with its own endpoint, not an Ollama model. But Ollaya ('Ollama for decision models') runs open Jev-like models (Laya, winnow, NLI) locally behind a TypeSafe-compatible API, so your Jev client works by changing one base-URL env var.
What is Ollaya?
A local runtime that pulls and serves open decision models by name, the way Ollama serves LLMs. It speaks TypeSafe's /v1/systemone format; GGUF models run on llama.cpp (CPU/CUDA/Metal) and on a Mac models like Laya run on the Apple GPU via MLX. It's a community project, not TypeSafe and not us.
Is a local model as good as hosted Jev?
Not out of the box. The open models trade accuracy for size and aren't RLCD-calibrated, so their confidence isn't something you can trust a production threshold against until you validate it on your own data. Hosted Jev is calibrated and pinned, which is the point of paying for it.
Which local model should I start with?
Laya (ModernBERT-large, 421M) is the usual starting point — it answers in a fraction of a second on CPU. If you have an NVIDIA GPU, GGUF models like winnow are worth trying. Validate whichever you pick on your own labelled data before wiring a threshold to it.
See also: Run a decision model locally · Laya vs Jev · Jev alternatives · Pricing
Want calibrated, with zero ops?
Local is great when you'll own the validation. For a calibrated decision with nothing to run, try hosted Jev free in the browser, then grab a jv_live_ key.