Jev engineering for coding agents
A coding agent spends most of its tokens on small judgment calls between edits — which skill to load, whether a command is safe to run, whether the task is done. Jev engineering is the practice of pulling those bounded calls out of the big model and handing them to Jev: one typed, calibrated decision in about 70–500ms instead of another paragraph to generate and parse.
The model behind Claude Code, Cursor or Codex is excellent at planning and writing code and expensive at everything else. Every loop it also makes dozens of routine choices — pick a tool, judge a diff, decide whether to keep going — and each one is a fresh generation you wait on and parse. Jev engineering moves those choices to a model built only for them: you send the state plus a typed question, and get back a value your code branches on, with a calibrated confidence.
What a coding agent should offload to Jev
Not everything — only the calls with a finite answer space. If the step produces code or prose, it stays with the big model. If it picks from a fixed set, judges a yes/no, or rates something on a scale, it's a Jev call:
| Decision in the loop | Primitive | What comes back |
|---|---|---|
| Which tool or skill to invoke next | choice | one key from your fixed set |
| Is this shell command / edit safe to run unattended? | noul | a calibrated yes-probability |
| Is the command destructive (delete / force-push / deploy / migrate)? | choice | a risk class to hard-gate on |
| Did this step actually succeed? | noul | a calibrated yes-probability |
| Keep iterating or stop and ask the human? | noul | a calibrated continue-probability |
| Which context chunks are worth keeping on compaction? | score | a 0–1 relevance score per chunk |
Each row is one question. You can send several in the same request and they're all answered in a single parallel pass — so gating a command, classifying its risk and checking whether you're done can be one round trip, not three.
Drop it into your agent — no code
The fastest way to start is to give your agent the decide tool once, then paste a rule that makes it check risky commands before running them. Add the MCP server (set JEV_API_KEY=jv_live_...):
npx github:codaaiteam/jev-mcp # gives the agent a `decide` toolClaude Code — save this as .claude/skills/jev-gate/SKILL.md:
---
name: jev-gate
description: Gate risky, irreversible shell/git commands with Jev before running them; stop for the human when unsafe or destructive.
---
# Jev safety gate for coding agents
Before running any irreversible command — rm, git push --force, a deploy,
a DB migration, or anything that writes outside the repo — call the `decide`
tool (from jev-mcp) with:
state: the exact command + one line of why
questions:
safe: { type: "noul", instructions: "Safe and reversible to run without a human?" }
kind: { type: "choice", criteria: { read, write, destructive } }
If safe < 0.8 or kind == "destructive" -> STOP and ask the human. Else run it.
Read-only commands (ls, cat, git status, grep, tests) skip the check.Codex / opencode — paste the same rule into your AGENTS.md. The same pattern works for a deterministic PreToolUse hook the agent can't skip.
Or wire it yourself (code)
Two calls a coding agent makes constantly — which tool to reach for, and whether it's done — in a single request:
// One decide call, two typed answers, before the next step
const res = await fetch("https://jevtypesafeai.com/api/v1/decide", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: `Bearer ${process.env.JEV_KEY}`, // jv_live_...
},
body: JSON.stringify({
state: { task, lastOutput, filesTouched },
questions: {
tool: { type: "choice", instructions: "Which tool should run next?",
criteria: { edit: "change a file", run: "run a command/tests",
search: "search the codebase", done: "nothing left to do" } },
done: { type: "noul", instructions: "Is the task fully complete and verified?" },
},
}),
});
const { answers } = await res.json();
if (answers.done.noul > 0.85) return finish(); // confident it's done
await dispatch(answers.tool.choice); // else run the chosen toolanswers.tool.choice is locked to the keys you listed, so there is no invalid tool to handle; answers.done.noul is a calibrated probability RLCD training makes trustworthy. Tune the 0.85 threshold against your own run logs.
What Jev does not do
Jev is not a replacement for the model behind your coding agent. It doesn't write code, stream text, call tools or edit files — it returns one typed value per question. Keep the frontier model for generation and reasoning, and use Jev for the bounded calls around it. "Can't hallucinate" only means it can't return a value outside your schema; a decision can still be wrong, so keep a human on genuinely destructive or consequential actions regardless of confidence.
Which primitive to use
- choice — pick one tool, skill or route from a fixed set; or classify a command's risk
- noul — a calibrated yes/no such as 'safe to run?', 'did this step succeed?', 'done yet?'
- score — rate relevance or risk on a 0–1 scale when you want a graded threshold (e.g. context compaction)
Want working references? jevgrep uses a score call to rank code chunks, Jev for computer use gates clicks, and Jev for compaction scores what to keep when the context bloats — all the same decide call, different questions.
FAQ
Does Jev replace the model behind Claude Code or Cursor?
No. Jev doesn't write code, stream text, call tools or edit files. It's the decision layer for the bounded calls around the coding model — which tool to use, whether a command is safe, whether a step is done — returned as typed values your harness branches on.
Won't a decide call on every step slow the agent down?
Not meaningfully. A Jev decision returns in about 70–500ms in a single pass, versus seconds for a second LLM prompt, and you can batch several questions into one round trip — so vetting each step costs far less than the generation it guards.
How do I keep a destructive command from auto-running?
Add a choice question that classifies the command (read / write / destructive) and never auto-run the 'destructive' class — route it to a human regardless of confidence, or require a much higher noul threshold on top.
Do I need to change my agent framework to use this?
No. You can start with zero code by giving the agent the jev-mcp decide tool plus a skill or AGENTS.md rule, or call the /api/v1/decide endpoint directly from your own loop. Either way the coding model stays exactly as it is.
See also: Jev for computer use · jevgrep — ranked code search · Jev for context compaction · Jev MCP server · Jev as a judge · jev-mcp (GitHub)
Offload your coding agent's routine calls to Jev
Grab a jv_live_ key and turn the bounded calls in your loop into typed, calibrated decisions.