Model routing
Let Jev pick which model should handle a request.
jev-router is an open-source, OpenAI-compatible LLM router built on LiteLLM: clients send one model id ("jev-router") and Jev picks which model actually serves each request. A pre-call hook summarizes the messages, filters candidates by capability (image support, output length), and lets Jev choose — with a rules-based cheapest-eligible fallback when no key is set. It's the exact pattern below, in production: a compact typed decision, easy to threshold and audit, instead of embedding similarity or hand-tuned rules.
Experimenta ao vivo
Isto é a coisa real, não uma maqueta. Edita a entrada, carrega em Executar e o Jev devolve cada resposta tipada numa só ida e volta — grátis, sem registo. Agora imagina a mesma chamada disparada sobre milhares de itens em paralelo.
Pick a demo, tweak the input, and hit Run Jev.
As decisões que o Jev toma
Numa só chamada, o Jev avalia cada uma destas — em paralelo, sobre a mesma entrada:
How complex is this request?
pontua-a numa escala ordenada:
- trivial
- simple
- moderate
- hard
- very hard / architectural
Which model should handle it?
escolhe uma destas opções:
fast— fast cheap model — trivial edits, formatting, lookupsbalanced— mid model — normal features and fixesstrong— frontier model — architecture, migrations, hard reasoning
Will this likely require tool calls (reading files, running commands)?
devolve uma probabilidade calibrada de sim/não.
O pedido exato
Este é o payload real por trás da demo ao vivo — copia-o, muda o state e já estás a construir:
{
"model": "jev-latest",
"state": "Incoming user request to an AI coding agent: \"Refactor our auth service to support multi-tenant SSO with SAML and SCIM, keep backward compatibility, and write a migration plan.\" Available models: fast (cheap, small), balanced (mid), strong (expensive, frontier).",
"questions": {
"complexity": {
"type": "score",
"instructions": "How complex is this request?",
"criteria": [
"trivial",
"simple",
"moderate",
"hard",
"very hard / architectural"
]
},
"route": {
"type": "choice",
"instructions": "Which model should handle it?",
"criteria": {
"fast": "fast cheap model — trivial edits, formatting, lookups",
"balanced": "mid model — normal features and fixes",
"strong": "frontier model — architecture, migrations, hard reasoning"
}
},
"needs_tools": {
"type": "noul",
"instructions": "Will this likely require tool calls (reading files, running commands)?"
}
}
}Integra-o no teu código
Lê as respostas tipadas e ramifica em código simples — sem análise. Trata automaticamente os casos de alta confiança e encaminha os incertos para um modelo maior ou uma pessoa. É uma só chamada à API e a saída é grátis, por isso faz todas as perguntas que precisas de uma vez.
Constrói o teu
Cada cenário acima é uma só chamada à API. Experimenta qualquer um deles grátis no playground e depois obtém uma chave alojada para o lançar em minutos.