Task-completion check
Before an agent stops, ask Jev whether the work is actually finished and the claims hold up.
Agents love to declare victory early. "Diff verification" and "agent output checks" in the awesome-jev catalog cover the stop decision: before an agent ends its turn, verify that the task is genuinely done and that its claims are backed by evidence (tests actually passed, the change addresses the ask). Jev's statement-verification shape — a calibrated yes/no against the task, the diff and the test output — is exactly the primitive TypeSafe's own docs recommend for "checking whether a statement is true of a record before taking an action." Below a threshold, the agent keeps working instead of handing back a half-finished job.
Попробуйте вживую
Это настоящее, а не макет. Отредактируйте ввод, нажмите «Запустить», и Jev вернёт каждый типизированный ответ за один round trip — бесплатно, без регистрации. А теперь представьте тот же вызов, запущенный по тысячам элементов параллельно.
Pick a demo, tweak the input, and hit Run Jev.
Решения, которые принимает Jev
За один вызов Jev оценивает каждое из них — параллельно, на одном и том же вводе:
Does the change actually address the stated task?
возвращает откалиброванную вероятность да/нет.
Does the test output provide real evidence the fix works (not just unrelated passes)?
возвращает откалиброванную вероятность да/нет.
How confident should the agent be that it's safe to stop and hand back?
оценивает это по упорядоченной шкале:
- keep working
- borderline
- likely done
- clearly done
Точный запрос
Это реальная полезная нагрузка за живым демо — скопируйте её, измените state, и вы уже строите:
{
"model": "jev-latest",
"state": "Task: \"Fix the bug where refunds over the order total are silently accepted.\"\n\nAgent's final summary: \"Added a balance check in process_refund so over-total refunds now raise RefundError.\"\n\nDiff: added `if amount > order.remaining_balance: raise RefundError(...)` before the gateway call.\n\nTest output: `test_refund_over_total PASSED · test_refund_partial PASSED · 2 passed, 0 failed`",
"questions": {
"complete": {
"type": "noul",
"instructions": "Does the change actually address the stated task?"
},
"tests_back_it": {
"type": "noul",
"instructions": "Does the test output provide real evidence the fix works (not just unrelated passes)?"
},
"confidence": {
"type": "score",
"instructions": "How confident should the agent be that it's safe to stop and hand back?",
"criteria": [
"keep working",
"borderline",
"likely done",
"clearly done"
]
}
}
}Встройте это в свой код
Читайте типизированные ответы и ветвитесь обычным кодом — без разбора. Автоматически обрабатывайте случаи с высокой уверенностью и направляйте неопределённые к более крупной модели или человеку. Это один вызов API, а вывод бесплатен, так что задавайте сразу все нужные вопросы.
Постройте своё
Каждый сценарий выше — это один вызов API. Попробуйте любой из них бесплатно в playground, затем получите размещённый ключ, чтобы выпустить его за минуты.