Task-completion check
Before an agent stops, ask Jev whether the work is actually finished and the claims hold up.
Agents love to declare victory early. "Diff verification" and "agent output checks" in the awesome-jev catalog cover the stop decision: before an agent ends its turn, verify that the task is genuinely done and that its claims are backed by evidence (tests actually passed, the change addresses the ask). Jev's statement-verification shape — a calibrated yes/no against the task, the diff and the test output — is exactly the primitive TypeSafe's own docs recommend for "checking whether a statement is true of a record before taking an action." Below a threshold, the agent keeps working instead of handing back a half-finished job.
ライブで試す
これはモックではなく本物です。入力を編集して「実行」を押すと、Jev が 1 回の往復ですべての型付き回答を返します — 無料、登録不要。同じ呼び出しを数千件のアイテムに並列で実行する様子を想像してみてください。
デモを選んで入力を調整し、こちらを押してください Jev を実行.
Jev が下す意思決定
1 回の呼び出しで、Jev はこれらのそれぞれを — 同じ入力に対して並列に — 評価します:
Does the change actually address the stated task?
較正済みの yes/no の確率を返します。
Does the test output provide real evidence the fix works (not just unrelated passes)?
較正済みの yes/no の確率を返します。
How confident should the agent be that it's safe to stop and hand back?
順序付きの尺度で評価します:
- keep working
- borderline
- likely done
- clearly done
実際のリクエスト
これがライブデモの背後にある本物のペイロードです — コピーして、state を変えれば、もう構築できています:
{
"model": "jev-latest",
"state": "Task: \"Fix the bug where refunds over the order total are silently accepted.\"\n\nAgent's final summary: \"Added a balance check in process_refund so over-total refunds now raise RefundError.\"\n\nDiff: added `if amount > order.remaining_balance: raise RefundError(...)` before the gateway call.\n\nTest output: `test_refund_over_total PASSED · test_refund_partial PASSED · 2 passed, 0 failed`",
"questions": {
"complete": {
"type": "noul",
"instructions": "Does the change actually address the stated task?"
},
"tests_back_it": {
"type": "noul",
"instructions": "Does the test output provide real evidence the fix works (not just unrelated passes)?"
},
"confidence": {
"type": "score",
"instructions": "How confident should the agent be that it's safe to stop and hand back?",
"criteria": [
"keep working",
"borderline",
"likely done",
"clearly done"
]
}
}
}コードに組み込む
型付きの答えを読んで、素のコードで分岐します — 解析は不要です。確信度の高いケースは自動処理し、不確かなものはより大きなモデルや人へ振り分けます。1 回の API 呼び出しで出力は無料なので、必要な質問はすべて一度に投げてください。
自分のものを作る
上のどのシナリオも 1 回の API 呼び出しです。プレイグラウンドでどれでも無料で試してから、ホスト型キーを取得して数分でそのまま本番投入できます。