Task-completion check
Before an agent stops, ask Jev whether the work is actually finished and the claims hold up.
Agents love to declare victory early. "Diff verification" and "agent output checks" in the awesome-jev catalog cover the stop decision: before an agent ends its turn, verify that the task is genuinely done and that its claims are backed by evidence (tests actually passed, the change addresses the ask). Jev's statement-verification shape — a calibrated yes/no against the task, the diff and the test output — is exactly the primitive TypeSafe's own docs recommend for "checking whether a statement is true of a record before taking an action." Below a threshold, the agent keeps working instead of handing back a half-finished job.
Experimenta ao vivo
Isto é a coisa real, não uma maqueta. Edita a entrada, carrega em Executar e o Jev devolve cada resposta tipada numa só ida e volta — grátis, sem registo. Agora imagina a mesma chamada disparada sobre milhares de itens em paralelo.
Pick a demo, tweak the input, and hit Run Jev.
As decisões que o Jev toma
Numa só chamada, o Jev avalia cada uma destas — em paralelo, sobre a mesma entrada:
Does the change actually address the stated task?
devolve uma probabilidade calibrada de sim/não.
Does the test output provide real evidence the fix works (not just unrelated passes)?
devolve uma probabilidade calibrada de sim/não.
How confident should the agent be that it's safe to stop and hand back?
pontua-a numa escala ordenada:
- keep working
- borderline
- likely done
- clearly done
O pedido exato
Este é o payload real por trás da demo ao vivo — copia-o, muda o state e já estás a construir:
{
"model": "jev-latest",
"state": "Task: \"Fix the bug where refunds over the order total are silently accepted.\"\n\nAgent's final summary: \"Added a balance check in process_refund so over-total refunds now raise RefundError.\"\n\nDiff: added `if amount > order.remaining_balance: raise RefundError(...)` before the gateway call.\n\nTest output: `test_refund_over_total PASSED · test_refund_partial PASSED · 2 passed, 0 failed`",
"questions": {
"complete": {
"type": "noul",
"instructions": "Does the change actually address the stated task?"
},
"tests_back_it": {
"type": "noul",
"instructions": "Does the test output provide real evidence the fix works (not just unrelated passes)?"
},
"confidence": {
"type": "score",
"instructions": "How confident should the agent be that it's safe to stop and hand back?",
"criteria": [
"keep working",
"borderline",
"likely done",
"clearly done"
]
}
}
}Integra-o no teu código
Lê as respostas tipadas e ramifica em código simples — sem análise. Trata automaticamente os casos de alta confiança e encaminha os incertos para um modelo maior ou uma pessoa. É uma só chamada à API e a saída é grátis, por isso faz todas as perguntas que precisas de uma vez.
Constrói o teu
Cada cenário acima é uma só chamada à API. Experimenta qualquer um deles grátis no playground e depois obtém uma chave alojada para o lançar em minutos.