← Alle Anwendungsfälle

Task-completion check

Before an agent stops, ask Jev whether the work is actually finished and the claims hold up.

Coding-agent use case · output & completion checksawesome-jev-by-typesafe

Agents love to declare victory early. "Diff verification" and "agent output checks" in the awesome-jev catalog cover the stop decision: before an agent ends its turn, verify that the task is genuinely done and that its claims are backed by evidence (tests actually passed, the change addresses the ask). Jev's statement-verification shape — a calibrated yes/no against the task, the diff and the test output — is exactly the primitive TypeSafe's own docs recommend for "checking whether a statement is true of a record before taking an action." Below a threshold, the agent keeps working instead of handing back a half-finished job.

Live ausprobieren

Das ist echt, kein Mockup. Bearbeite die Eingabe, klicke auf Ausführen, und Jev liefert jede typisierte Antwort in einem einzigen Roundtrip — kostenlos, ohne Anmeldung. Stell dir nun denselben Aufruf vor, parallel über Tausende von Elementen ausgelöst.

POST jevtypesafeai.com/api/v1/decide
Status — die Eingabe, die Software an Jev übergibt384c
Fragen — die typisierten Entscheidungen, die du zurückbekommst
noulcomplete
Does the change actually address the stated task?
→ Wahrscheinlichkeit 0,0 … 1,0
noultests_back_it
Does the test output provide real evidence the fix works (not just unrelated passes)?
→ Wahrscheinlichkeit 0,0 … 1,0
scoreconfidence
How confident should the agent be that it's safe to stop and hand back?
→ 0…3 · 4 Stufen
echte API · kostenlos · ohne Anmeldung
Typisierte, kalibrierte Ausgabe erscheint hier.
Wähle ein Demo, passe die Eingabe an und klicke auf Jev ausführen.
API-Key holen →← Alle Anwendungsfälle

Die Entscheidungen, die Jev trifft

In einem einzigen Aufruf wertet Jev jede davon aus — parallel, gegen dieselbe Eingabe:

noulcomplete

Does the change actually address the stated task?

gibt eine kalibrierte Ja/Nein-Wahrscheinlichkeit zurück.

noultests_back_it

Does the test output provide real evidence the fix works (not just unrelated passes)?

gibt eine kalibrierte Ja/Nein-Wahrscheinlichkeit zurück.

scoreconfidence

How confident should the agent be that it's safe to stop and hand back?

bewertet es auf einer geordneten Skala:

  1. keep working
  2. borderline
  3. likely done
  4. clearly done

Die exakte Anfrage

Dies ist der echte Payload hinter der Live-Demo — kopiere ihn, ändere den state, und du baust schon:

{
  "model": "jev-latest",
  "state": "Task: \"Fix the bug where refunds over the order total are silently accepted.\"\n\nAgent's final summary: \"Added a balance check in process_refund so over-total refunds now raise RefundError.\"\n\nDiff: added `if amount > order.remaining_balance: raise RefundError(...)` before the gateway call.\n\nTest output: `test_refund_over_total PASSED · test_refund_partial PASSED · 2 passed, 0 failed`",
  "questions": {
    "complete": {
      "type": "noul",
      "instructions": "Does the change actually address the stated task?"
    },
    "tests_back_it": {
      "type": "noul",
      "instructions": "Does the test output provide real evidence the fix works (not just unrelated passes)?"
    },
    "confidence": {
      "type": "score",
      "instructions": "How confident should the agent be that it's safe to stop and hand back?",
      "criteria": [
        "keep working",
        "borderline",
        "likely done",
        "clearly done"
      ]
    }
  }
}

Binde es in deinen Code ein

Lies die typisierten Antworten und verzweige in einfachem Code — ohne Parsing. Bearbeite die hochkonfidenten Fälle automatisch und leite die unsicheren an ein größeres Modell oder einen Menschen weiter. Es ist ein einziger API-Aufruf und Output ist kostenlos, stelle also alle Fragen, die du brauchst, auf einmal.

Bau dein eigenes

Jedes Szenario oben ist ein einziger API-Aufruf. Teste jedes davon kostenlos im Playground und hol dir dann einen gehosteten Key, um es in Minuten auszuliefern.

Diese Demo ausführen ▶API-Key holen →
Task-completion check — ein Jev-Anwendungsfall mit einer Live-Demo · Jev by TypeSafe AI