Task-completion check
Before an agent stops, ask Jev whether the work is actually finished and the claims hold up.
Agents love to declare victory early. "Diff verification" and "agent output checks" in the awesome-jev catalog cover the stop decision: before an agent ends its turn, verify that the task is genuinely done and that its claims are backed by evidence (tests actually passed, the change addresses the ask). Jev's statement-verification shape — a calibrated yes/no against the task, the diff and the test output — is exactly the primitive TypeSafe's own docs recommend for "checking whether a statement is true of a record before taking an action." Below a threshold, the agent keeps working instead of handing back a half-finished job.
Canlı deneyin
Bu gerçek olanı, bir maket değil. Girdiyi düzenleyin, Çalıştır'a basın ve Jev her tipli cevabı tek bir gidiş-dönüşte döndürsün — ücretsiz, kayıt yok. Şimdi aynı çağrının binlerce öğe üzerinde paralel olarak tetiklendiğini hayal edin.
Pick a demo, tweak the input, and hit Run Jev.
Jev'in verdiği kararlar
Tek bir çağrıda Jev bunların her birini değerlendirir — paralel olarak, aynı girdiye karşı:
Does the change actually address the stated task?
kalibre edilmiş bir evet/hayır olasılığı döndürür.
Does the test output provide real evidence the fix works (not just unrelated passes)?
kalibre edilmiş bir evet/hayır olasılığı döndürür.
How confident should the agent be that it's safe to stop and hand back?
onu sıralı bir ölçekte puanlar:
- keep working
- borderline
- likely done
- clearly done
Tam istek
Bu, canlı demonun arkasındaki gerçek payload — kopyalayın, state'i değiştirin ve inşa etmeye başlayın:
{
"model": "jev-latest",
"state": "Task: \"Fix the bug where refunds over the order total are silently accepted.\"\n\nAgent's final summary: \"Added a balance check in process_refund so over-total refunds now raise RefundError.\"\n\nDiff: added `if amount > order.remaining_balance: raise RefundError(...)` before the gateway call.\n\nTest output: `test_refund_over_total PASSED · test_refund_partial PASSED · 2 passed, 0 failed`",
"questions": {
"complete": {
"type": "noul",
"instructions": "Does the change actually address the stated task?"
},
"tests_back_it": {
"type": "noul",
"instructions": "Does the test output provide real evidence the fix works (not just unrelated passes)?"
},
"confidence": {
"type": "score",
"instructions": "How confident should the agent be that it's safe to stop and hand back?",
"criteria": [
"keep working",
"borderline",
"likely done",
"clearly done"
]
}
}
}Kodunuza entegre edin
Tipli cevapları okuyun ve düz kodda dallanın — ayrıştırma yok. Yüksek güvenli durumları otomatik ele alın ve belirsiz olanları daha büyük bir modele veya bir insana yönlendirin. Tek bir API çağrısıdır ve çıktı ücretsizdir, bu yüzden ihtiyaç duyduğunuz her soruyu aynı anda sorun.
Kendinizinkini inşa edin
Yukarıdaki her senaryo tek bir API çağrısıdır. Herhangi birini playground'da ücretsiz deneyin, ardından onu dakikalar içinde yayınlamak için barındırılan bir anahtar edinin.