Using Jev in a RAG pipeline
Retrieval gives you candidate chunks ranked by embedding similarity — which isn't the same as relevance. Jev is a fast, typed second pass: for each chunk it returns a calibrated 'does this actually answer the question?' in roughly 70–500ms, so you drop the near-misses before they waste context and dilute the answer.
The weak link in most RAG stacks is the gap between 'similar' and 'useful'. Vector search is recall-oriented, so the top-k is full of chunks that are topically close but don't answer the question. Jev is TypeSafe AI's System One model: hand it the query and a chunk as state, ask a typed question, and it emits a yes/no or a score with calibrated confidence in a single non-autoregressive pass — no prose to parse, no invalid output, cheap enough to run on every candidate.
Where Jev fits in the pipeline
- After retrieval, before the LLM — rerank or filter the top-k so only genuinely relevant chunks spend context budget
- As a relevance gate — a noul 'does this chunk answer the question?' per chunk, keep those above your threshold
- As a grounding check — after generation, ask whether the answer is supported by the retrieved chunks before you return it
- As a query router — a choice over indexes/datasources so you only search the ones that matter
Filter the top-k in code
// Rerank/filter retrieved chunks with a typed relevance call per chunk.
async function keepRelevant(query, chunks) {
const judged = await Promise.all(chunks.map(async (chunk) => {
const res = await fetch("https://jevtypesafeai.com/api/v1/decide", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.JEV_KEY}`, // jv_live_...
},
body: JSON.stringify({
state: { query, chunk: chunk.text },
questions: {
relevant: { type: "noul", instructions: "Does this passage help answer the query?" },
usefulness: { type: "score", instructions: "How useful is it?", criteria: ["none", "weak", "strong"] },
},
}),
});
const { answers } = await res.json();
return { chunk, keep: answers.relevant.noul, rank: answers.usefulness.score };
}));
return judged
.filter((j) => j.keep > 0.6) // calibrated relevance gate
.sort((a, b) => b.rank - a.rank) // reorder by usefulness
.map((j) => j.chunk); // feed only these to the LLM
}answers.relevant.noul is a calibrated 0–1 relevance probability (RLCD training makes the number trustworthy), and answers.usefulness.score is locked to your ordered scale — so you gate on one and rerank on the other, with nothing to parse. Running this on 20 candidates adds a few hundred milliseconds and a fraction of a cent, and typically lets you cut the context you send the LLM by more than half.
Why a typed pass beats prompting an LLM to rerank
| LLM reranker | Jev relevance pass | |
|---|---|---|
| Output | Text/JSON to parse | Typed yes/no + score |
| Invalid output | Possible | Impossible (schema-locked) |
| Confidence | You estimate it | Calibrated, returned |
| Latency per chunk | Seconds | ~70–500ms |
| Cost at top-k | Frontier-LLM pricing | Far cheaper (output free) |
You can also ground the final answer the same way: send the generated answer plus the retrieved chunks as state and ask a noul 'is this answer fully supported by these passages?' — a cheap, typed hallucination check before you return anything to the user.
See the primitive running
Semantic Grep filters lines by meaning rather than regex, and Jev as a judge covers the same calibrated-verdict pattern end to end. Point your own RAG pipeline at POST https://jevtypesafeai.com/api/v1/decide with a jv_live_ key (official upstream POST https://api.typesafe.ai/v1/systemone) and you get the same typed relevance call.
FAQ
How does Jev help a RAG pipeline?
It closes the gap between 'similar' and 'relevant'. After vector search returns the top-k, Jev makes a fast, typed call per chunk — a calibrated 'does this answer the question?' yes/no and an optional usefulness score — so you filter and rerank before the LLM, sending it only the chunks that matter instead of everything that embedded closely.
Isn't a call per chunk too slow or expensive?
No — each Jev decision returns in roughly 70–500ms in a single pass, and you run the chunks in parallel, so reranking 20 candidates adds a few hundred milliseconds and a fraction of a cent (output tokens are free). The context you save usually makes the overall call cheaper, not more expensive.
Can Jev do the embeddings or the generation too?
No — Jev is a decision model, not an embedder or a text generator. Keep your existing retriever and your existing LLM; Jev slots in between them as the typed relevance filter, and optionally after generation as a grounding check.
How do I pick the relevance threshold?
Start around 0.6 on the calibrated noul and tune against a small labelled set from your own corpus. Because the probability is RLCD-calibrated, the threshold behaves predictably — raise it to send the LLM less, lower it if you're dropping useful chunks.
See also: Jev as a judge · Semantic Grep · Jev classifier · How RLCD works · Playground
Make your RAG retrieve relevance, not just similarity
Grab a jv_live_ key and add a typed, calibrated relevance pass between retrieval and your LLM.