← Use cases

Jev's context window

Jev's context window is the budget for one decide call: roughly 64k tokens for your state plus all the questions you ask, with a tighter ~32k ceiling on the state plus any single question. Because Jev returns a small typed value rather than generated text, there's no output side of the budget to manage — the whole window is yours for context.

Jev is TypeSafe AI's System One model: you send it state (the text or JSON you want judged) and one or more typed questions, and it returns structured values — choice, score, noul — in a single non-autoregressive parallel pass. The context window is simply how much of that state-plus-questions payload fits in one call.

~64k tokens — state + all questions state questions ~32k cap — state + the single longest question state output ~0 Output is a small typed value, not generated tokens — it doesn't consume the window. Counts are estimates; the model tokenizes.
Two budgets: ~64k tokens for state plus all questions, and a tighter ~32k ceiling for state plus the single longest question. Output is a typed value, so it doesn't eat the window.

The two budgets

These are token counts, and exact tokenization is done by the model — any figure you compute client-side from characters is an estimate. As a rough guide, clean English prose runs around 4–6 characters per token, so a ~64k budget is on the order of a few hundred KB of text, not a few MB.

How it differs from an LLM context window

With a chat LLM, the context window is shared between your prompt and the tokens the model generates back — a long answer eats into what you could have spent on context, and you pay for both. Jev only emits a compact typed value, so the entire window goes to the decision's input. You size the window around how much state the judgment actually needs, nothing else.

If you hit the limit

An over-budget request comes back as a max_tokens_exceeded error rather than a silently truncated decision — Jev won't quietly drop half your state and judge the rest. The fix is almost always to trim the state to only the context the decision needs: one ticket, not the whole thread; the diff, not the entire file; the five candidate rows, not the table.

// A decision rarely needs the whole document — pass the span that matters.
const res = await fetch("https://jevtypesafeai.com/api/v1/decide", {
  method: "POST",
  headers: { "Content-Type": "application/json", Authorization: `Bearer ${process.env.JEV_API_KEY}` },
  body: JSON.stringify({
    state: { ticket: ticket.body },           // just this ticket, not the inbox
    questions: {
      urgent: { type: "noul", instructions: "Is this urgent?" },
      queue: { type: "choice", instructions: "Route it.",
               criteria: { billing: "payments", bug: "something broken", sales: "pricing" } },
    },
  }),
});

If the state genuinely is large — a long document you need judged as a whole — score it in chunks and combine the results, or use Jev to compact the context first: score each section for relevance and keep only the parts that matter before the real decision. That keeps every call comfortably inside the window.

Picking what goes in the window

FAQ

What is Jev's context window?

Roughly 64k tokens for a whole decide call — your state plus all the questions — with a tighter ~32k ceiling on the state plus any single question. These are estimates; the model does the exact tokenization.

Does Jev's output use up the context window?

No. Jev returns a small typed value (a choice, score, or noul with calibrated confidence), not a generated completion, so there's no output-token budget. Unlike an LLM, the whole window is available for your input state and questions.

What happens if my request is too big?

You get a max_tokens_exceeded error instead of a silently truncated decision. Trim the state to just the context the judgment needs, split a large document into scored chunks, or compact the context with Jev first, then retry.

How do I count tokens for Jev?

Exact tokenization is performed by the model, so any count you compute yourself is an estimate. A rough rule for clean prose is 4–6 characters per token — enough to sanity-check that a payload is nowhere near the ~64k budget before you send it.

How is this different from an LLM's context window?

An LLM splits its window between your prompt and the text it generates, and you pay for both. Jev doesn't generate text, so the entire window is for context and only the input is metered — you size it purely around how much state the decision needs.

See also: Compact context with Jev · How Jev works · API reference · Docs

Size your first decision

Grab a jv_live_ key and send a decide call — trim the state to what the judgment needs and you'll stay well inside the window.

▶ Try Jev freeGet an API key →
Jev context window — the token budget for a decision · Jev by TypeSafe AI