Leaderboard
JevServe-bench evaluates System One decision models across LLM system decisions, including routing, caching, scheduling, verification, guardrails and tool use.
Tasks
Each task is one decision a serving system makes, with fixed questions and a fixed downstream rule, so the decision model is the only thing that changes between entries. On every task, 0 is what the system does without the model and 100 is an oracle that knows the outcome in advance.
Methodology
Every task scores what a serving component achieves when a decision model makes its decision, on a scale from the system's default (0) to an oracle that knows the outcome in advance (100).
A score is how much of the gap between the serving default and the oracle the component closes. Below 0, it does worse than the default.
Classification tasks use AUROC, and the thinking and cascade tasks act on requests in the order the model ranks them, so a model is judged on its ordering rather than on a tuned cutoff.
The average is the plain mean of the task scores. A category's score is the mean of its tasks; the average is not a mean of the categories.
A 95% interval: a bootstrap over the task's items, a prompt's responses resampled together, or for routing and scheduling the spread over data splits and arrival draws. Ranks and bold scores compare models on the same resamples, paired, and call two models tied unless one is ahead beyond the 95% interval of their difference; tied models share a rank.
p50 is the decision model's median single-request latency on the submitter's setup, measured after a warm-up, so it is reported rather than compared. It enters one score: in the scheduling task every request waits that long for its prediction.
A request that still fails after retries is scored as unanswered, with uninformative probabilities. It is never dropped.
The benchmark code is not public yet. Any decision model that speaks the interface below can take part once it is; this section will then list the commands to run it and submit.
POST /v1/systemoneAny HTTP server that speaks TypeSafe's System One API works, which hosted Jev and the Jev-compatible open models already do. Each request carries the text and the questions; the reply carries probabilities.
# an example exchange POST /v1/systemone {"state": "Prove that there are infinitely many primes.", "questions": {"needs_thinking": {"type": "noul", "instructions": "Does answering this request correctly require thinking step by step at length before responding, rather than answering directly?"}}} # the reply: P(yes) for a yes/no question {"answers": {"needs_thinking": {"noul": 0.91}}}
Prepare the data first (README, “Data”). A full run asks every question once and caches the responses,
so an interrupted run resumes; --limit 50 checks the whole pipeline on 50 items per task.
git clone && cd jevserve-bench uv venv .venv && uv pip install --python .venv/bin/python -e ".[prepare]" .venv/bin/jevserve run --backend http --base-url http://127.0.0.1:8000 --model my-model --name my-model
Open a pull request with results/my-model/: scores.json,
environment.json and the answer files. We re-score the answers and add a model card before
listing it.