Task-completion check
Before an agent stops, ask Jev whether the work is actually finished and the claims hold up.
Agents love to declare victory early. "Diff verification" and "agent output checks" in the awesome-jev catalog cover the stop decision: before an agent ends its turn, verify that the task is genuinely done and that its claims are backed by evidence (tests actually passed, the change addresses the ask). Jev's statement-verification shape — a calibrated yes/no against the task, the diff and the test output — is exactly the primitive TypeSafe's own docs recommend for "checking whether a statement is true of a record before taking an action." Below a threshold, the agent keeps working instead of handing back a half-finished job.
Try it live
This is the real thing, not a mockup. Edit the input, hit Run, and Jev returns every typed answer in one round trip — free, no signup. Now picture the same call fired across thousands of items in parallel.
Pick a demo, tweak the input, and hit Run Jev.
The decisions Jev makes
In a single call, Jev evaluates each of these — in parallel, against the same input:
Does the change actually address the stated task?
returns a calibrated yes/no probability.
Does the test output provide real evidence the fix works (not just unrelated passes)?
returns a calibrated yes/no probability.
How confident should the agent be that it's safe to stop and hand back?
rates it on an ordered scale:
- keep working
- borderline
- likely done
- clearly done
The exact request
This is the real payload behind the live demo — copy it, change the state, and you're building:
{
"model": "jev-latest",
"state": "Task: \"Fix the bug where refunds over the order total are silently accepted.\"\n\nAgent's final summary: \"Added a balance check in process_refund so over-total refunds now raise RefundError.\"\n\nDiff: added `if amount > order.remaining_balance: raise RefundError(...)` before the gateway call.\n\nTest output: `test_refund_over_total PASSED · test_refund_partial PASSED · 2 passed, 0 failed`",
"questions": {
"complete": {
"type": "noul",
"instructions": "Does the change actually address the stated task?"
},
"tests_back_it": {
"type": "noul",
"instructions": "Does the test output provide real evidence the fix works (not just unrelated passes)?"
},
"confidence": {
"type": "score",
"instructions": "How confident should the agent be that it's safe to stop and hand back?",
"criteria": [
"keep working",
"borderline",
"likely done",
"clearly done"
]
}
}
}Wire it into your code
Read the typed answers and branch in plain code — no parsing. Auto-handle the high-confidence cases and route the uncertain ones to a bigger model or a human. It's one API call and output is free, so ask every question you need at once.
Build your own
Every scenario above is a single API call. Try any of them free in the playground, then get a hosted key to ship it in minutes.