← Use cases

Using Jev in a RAG pipeline

Retrieval gives you candidate chunks ranked by embedding similarity — which isn't the same as relevance. Jev is a fast, typed second pass: for each chunk it returns a calibrated 'does this actually answer the question?' in roughly 70–500ms, so you drop the near-misses before they waste context and dilute the answer.

The weak link in most RAG stacks is the gap between 'similar' and 'useful'. Vector search is recall-oriented, so the top-k is full of chunks that are topically close but don't answer the question. Jev is TypeSafe AI's System One model: hand it the query and a chunk as state, ask a typed question, and it emits a yes/no or a score with calibrated confidence in a single non-autoregressive pass — no prose to parse, no invalid output, cheap enough to run on every candidate.

Jev verdict confidence? calibrated high → auto-accept low → human review
Each retrieved chunk gets a typed relevance verdict with calibrated confidence: keep the confident-relevant ones, drop the rest before they reach the LLM.

Where Jev fits in the pipeline

Filter the top-k in code

// Rerank/filter retrieved chunks with a typed relevance call per chunk.
async function keepRelevant(query, chunks) {
  const judged = await Promise.all(chunks.map(async (chunk) => {
    const res = await fetch("https://jevtypesafeai.com/api/v1/decide", {
      method: "POST",
      headers: {
        "Content-Type": "application/json",
        Authorization: `Bearer ${process.env.JEV_KEY}`, // jv_live_...
      },
      body: JSON.stringify({
        state: { query, chunk: chunk.text },
        questions: {
          relevant: { type: "noul", instructions: "Does this passage help answer the query?" },
          usefulness: { type: "score", instructions: "How useful is it?", criteria: ["none", "weak", "strong"] },
        },
      }),
    });
    const { answers } = await res.json();
    return { chunk, keep: answers.relevant.noul, rank: answers.usefulness.score };
  }));

  return judged
    .filter((j) => j.keep > 0.6)              // calibrated relevance gate
    .sort((a, b) => b.rank - a.rank)          // reorder by usefulness
    .map((j) => j.chunk);                     // feed only these to the LLM
}

answers.relevant.noul is a calibrated 0–1 relevance probability (RLCD training makes the number trustworthy), and answers.usefulness.score is locked to your ordered scale — so you gate on one and rerank on the other, with nothing to parse. Running this on 20 candidates adds a few hundred milliseconds and a fraction of a cent, and typically lets you cut the context you send the LLM by more than half.

Why a typed pass beats prompting an LLM to rerank

LLM rerankerJev relevance pass
OutputText/JSON to parseTyped yes/no + score
Invalid outputPossibleImpossible (schema-locked)
ConfidenceYou estimate itCalibrated, returned
Latency per chunkSeconds~70–500ms
Cost at top-kFrontier-LLM pricingFar cheaper (output free)

You can also ground the final answer the same way: send the generated answer plus the retrieved chunks as state and ask a noul 'is this answer fully supported by these passages?' — a cheap, typed hallucination check before you return anything to the user.

See the primitive running

Semantic Grep filters lines by meaning rather than regex, and Jev as a judge covers the same calibrated-verdict pattern end to end. Point your own RAG pipeline at POST https://jevtypesafeai.com/api/v1/decide with a jv_live_ key (official upstream POST https://api.typesafe.ai/v1/systemone) and you get the same typed relevance call.

FAQ

How does Jev help a RAG pipeline?

It closes the gap between 'similar' and 'relevant'. After vector search returns the top-k, Jev makes a fast, typed call per chunk — a calibrated 'does this answer the question?' yes/no and an optional usefulness score — so you filter and rerank before the LLM, sending it only the chunks that matter instead of everything that embedded closely.

Isn't a call per chunk too slow or expensive?

No — each Jev decision returns in roughly 70–500ms in a single pass, and you run the chunks in parallel, so reranking 20 candidates adds a few hundred milliseconds and a fraction of a cent (output tokens are free). The context you save usually makes the overall call cheaper, not more expensive.

Can Jev do the embeddings or the generation too?

No — Jev is a decision model, not an embedder or a text generator. Keep your existing retriever and your existing LLM; Jev slots in between them as the typed relevance filter, and optionally after generation as a grounding check.

How do I pick the relevance threshold?

Start around 0.6 on the calibrated noul and tune against a small labelled set from your own corpus. Because the probability is RLCD-calibrated, the threshold behaves predictably — raise it to send the LLM less, lower it if you're dropping useful chunks.

See also: Jev as a judge · Semantic Grep · Jev classifier · How RLCD works · Playground

Make your RAG retrieve relevance, not just similarity

Grab a jv_live_ key and add a typed, calibrated relevance pass between retrieval and your LLM.

▶ Try Jev freeGet an API key →
Jev for RAG — typed relevance filtering and reranking · Jev by TypeSafe AI