Jev's context window
Jev's context window is the budget for one decide call: roughly 64k tokens for your state plus all the questions you ask, with a tighter ~32k ceiling on the state plus any single question. Because Jev returns a small typed value rather than generated text, there's no output side of the budget to manage — the whole window is yours for context.
Jev is TypeSafe AI's System One model: you send it state (the text or JSON you want judged) and one or more typed questions, and it returns structured values — choice, score, noul — in a single non-autoregressive parallel pass. The context window is simply how much of that state-plus-questions payload fits in one call.
The two budgets
- ~64k tokens for the whole request — your state plus every question you attach to the call
- ~32k tokens for the state plus the single longest question — the per-question ceiling bites first when one question carries a long rubric or criteria set
- Output isn't counted: unlike an LLM, Jev doesn't generate a long completion, so there's no output-token budget competing for the window
These are token counts, and exact tokenization is done by the model — any figure you compute client-side from characters is an estimate. As a rough guide, clean English prose runs around 4–6 characters per token, so a ~64k budget is on the order of a few hundred KB of text, not a few MB.
How it differs from an LLM context window
With a chat LLM, the context window is shared between your prompt and the tokens the model generates back — a long answer eats into what you could have spent on context, and you pay for both. Jev only emits a compact typed value, so the entire window goes to the decision's input. You size the window around how much state the judgment actually needs, nothing else.
If you hit the limit
An over-budget request comes back as a max_tokens_exceeded error rather than a silently truncated decision — Jev won't quietly drop half your state and judge the rest. The fix is almost always to trim the state to only the context the decision needs: one ticket, not the whole thread; the diff, not the entire file; the five candidate rows, not the table.
// A decision rarely needs the whole document — pass the span that matters.
const res = await fetch("https://jevtypesafeai.com/api/v1/decide", {
method: "POST",
headers: { "Content-Type": "application/json", Authorization: `Bearer ${process.env.JEV_API_KEY}` },
body: JSON.stringify({
state: { ticket: ticket.body }, // just this ticket, not the inbox
questions: {
urgent: { type: "noul", instructions: "Is this urgent?" },
queue: { type: "choice", instructions: "Route it.",
criteria: { billing: "payments", bug: "something broken", sales: "pricing" } },
},
}),
});If the state genuinely is large — a long document you need judged as a whole — score it in chunks and combine the results, or use Jev to compact the context first: score each section for relevance and keep only the parts that matter before the real decision. That keeps every call comfortably inside the window.
Picking what goes in the window
- Include only the state the decision depends on — extra context costs budget and can dilute the signal
- Keep question instructions and criteria tight; a sprawling rubric on one question is what trips the ~32k per-question ceiling
- For many items, prefer many small calls over one giant call — each decision is cheap and stays well under budget
FAQ
What is Jev's context window?
Roughly 64k tokens for a whole decide call — your state plus all the questions — with a tighter ~32k ceiling on the state plus any single question. These are estimates; the model does the exact tokenization.
Does Jev's output use up the context window?
No. Jev returns a small typed value (a choice, score, or noul with calibrated confidence), not a generated completion, so there's no output-token budget. Unlike an LLM, the whole window is available for your input state and questions.
What happens if my request is too big?
You get a max_tokens_exceeded error instead of a silently truncated decision. Trim the state to just the context the judgment needs, split a large document into scored chunks, or compact the context with Jev first, then retry.
How do I count tokens for Jev?
Exact tokenization is performed by the model, so any count you compute yourself is an estimate. A rough rule for clean prose is 4–6 characters per token — enough to sanity-check that a payload is nowhere near the ~64k budget before you send it.
How is this different from an LLM's context window?
An LLM splits its window between your prompt and the text it generates, and you pay for both. Jev doesn't generate text, so the entire window is for context and only the input is metered — you size it purely around how much state the decision needs.
See also: Compact context with Jev · How Jev works · API reference · Docs
Size your first decision
Grab a jv_live_ key and send a decide call — trim the state to what the judgment needs and you'll stay well inside the window.