Task-completion check
Before an agent stops, ask Jev whether the work is actually finished and the claims hold up.
Agents love to declare victory early. "Diff verification" and "agent output checks" in the awesome-jev catalog cover the stop decision: before an agent ends its turn, verify that the task is genuinely done and that its claims are backed by evidence (tests actually passed, the change addresses the ask). Jev's statement-verification shape — a calibrated yes/no against the task, the diff and the test output — is exactly the primitive TypeSafe's own docs recommend for "checking whether a statement is true of a record before taking an action." Below a threshold, the agent keeps working instead of handing back a half-finished job.
Pruébalo en vivo
Esto es lo real, no una maqueta. Edita la entrada, pulsa Ejecutar y Jev devuelve cada respuesta tipada en un solo viaje de ida y vuelta — gratis, sin registro. Ahora imagina la misma llamada disparada sobre miles de elementos en paralelo.
Elige una demo, ajusta la entrada y pulsa Ejecutar Jev.
Las decisiones que toma Jev
En una sola llamada, Jev evalúa cada una de estas — en paralelo, sobre la misma entrada:
Does the change actually address the stated task?
devuelve una probabilidad calibrada de sí/no.
Does the test output provide real evidence the fix works (not just unrelated passes)?
devuelve una probabilidad calibrada de sí/no.
How confident should the agent be that it's safe to stop and hand back?
la califica en una escala ordenada:
- keep working
- borderline
- likely done
- clearly done
La petición exacta
Este es el payload real detrás de la demo en vivo — cópialo, cambia el state y ya estás construyendo:
{
"model": "jev-latest",
"state": "Task: \"Fix the bug where refunds over the order total are silently accepted.\"\n\nAgent's final summary: \"Added a balance check in process_refund so over-total refunds now raise RefundError.\"\n\nDiff: added `if amount > order.remaining_balance: raise RefundError(...)` before the gateway call.\n\nTest output: `test_refund_over_total PASSED · test_refund_partial PASSED · 2 passed, 0 failed`",
"questions": {
"complete": {
"type": "noul",
"instructions": "Does the change actually address the stated task?"
},
"tests_back_it": {
"type": "noul",
"instructions": "Does the test output provide real evidence the fix works (not just unrelated passes)?"
},
"confidence": {
"type": "score",
"instructions": "How confident should the agent be that it's safe to stop and hand back?",
"criteria": [
"keep working",
"borderline",
"likely done",
"clearly done"
]
}
}
}Intégralo en tu código
Lee las respuestas tipadas y ramifica en código plano — sin parseo. Gestiona automáticamente los casos de alta confianza y deriva los inciertos a un modelo más grande o a una persona. Es una sola llamada a la API y la salida es gratis, así que haz todas las preguntas que necesites de una vez.
Construye el tuyo
Cada escenario de arriba es una sola llamada a la API. Prueba cualquiera de ellos gratis en el playground y luego obtén una clave alojada para lanzarlo en minutos.