Task-completion check
Before an agent stops, ask Jev whether the work is actually finished and the claims hold up.
Agents love to declare victory early. "Diff verification" and "agent output checks" in the awesome-jev catalog cover the stop decision: before an agent ends its turn, verify that the task is genuinely done and that its claims are backed by evidence (tests actually passed, the change addresses the ask). Jev's statement-verification shape — a calibrated yes/no against the task, the diff and the test output — is exactly the primitive TypeSafe's own docs recommend for "checking whether a statement is true of a record before taking an action." Below a threshold, the agent keeps working instead of handing back a half-finished job.
在线试用
这是真实的东西,不是模型演示。编辑输入、点击运行,Jev 会在一次往返中返回每一个类型化答案——免费、无需注册。现在想象同一次调用并行地跑在数千条数据上。
选一个示例,调整输入,然后点击 运行 Jev.
Jev 做出的决策
在一次调用中,Jev 针对同一份输入并行评估以下每一项:
Does the change actually address the stated task?
返回一个已校准的是/否概率。
Does the test output provide real evidence the fix works (not just unrelated passes)?
返回一个已校准的是/否概率。
How confident should the agent be that it's safe to stop and hand back?
在一个有序量表上给它打分:
- keep working
- borderline
- likely done
- clearly done
确切的请求
这就是实时演示背后真实的载荷——复制它,改一下 state,你就开始构建了:
{
"model": "jev-latest",
"state": "Task: \"Fix the bug where refunds over the order total are silently accepted.\"\n\nAgent's final summary: \"Added a balance check in process_refund so over-total refunds now raise RefundError.\"\n\nDiff: added `if amount > order.remaining_balance: raise RefundError(...)` before the gateway call.\n\nTest output: `test_refund_over_total PASSED · test_refund_partial PASSED · 2 passed, 0 failed`",
"questions": {
"complete": {
"type": "noul",
"instructions": "Does the change actually address the stated task?"
},
"tests_back_it": {
"type": "noul",
"instructions": "Does the test output provide real evidence the fix works (not just unrelated passes)?"
},
"confidence": {
"type": "score",
"instructions": "How confident should the agent be that it's safe to stop and hand back?",
"criteria": [
"keep working",
"borderline",
"likely done",
"clearly done"
]
}
}
}把它接进你的代码
读取类型化的答案,用普通代码分支判断——无需解析。自动处理高置信度的情形,把不确定的路由给更大的模型或人工。这只是一次 API 调用,而且输出免费,所以把你需要的每个问题一次都问了吧。