Archived board: v0.1 as published on 30 Sep 2026 · the current board
Leaderboard
A decision model reads a request and answers typed questions with probabilities; it writes no text. JevServe-bench puts one in charge of eight decisions an LLM serving system makes, from routing and scheduling to guardrails, and scores what each component then achieves. On every task, 0 is what the system does without the model and 100 is an oracle that knows the outcome in advance.
The average is the plain mean of the eight task scores. Select a column to sort by it, or a model's name for the raw numbers behind its scores.
p50 is the decision model's median single-request latency on the submitter's setup, so it is reported, not compared. It enters one score: in the scheduling task every request waits that long for its prediction.
Each task is one decision a serving system makes, with fixed questions and a fixed downstream rule, so the decision model is the only thing that changes between entries. In request order:
A score is how much of the gap between the serving default and the oracle the component closes. Below 0, it does worse than the default.
Classification tasks use AUROC, and the thinking and cascade tasks act on requests in the order the model ranks them, so a model is judged on its ordering rather than on a tuned cutoff.
Tier A replays logged outcomes and human labels, so the same answers score the same on any machine. Tier B runs a vLLM scheduler simulator calibrated on an H200 serving Qwen3-8B, within 2–4% of the real server. The measured decision latency is the only hardware-dependent input.
A request that still fails after retries is scored as unanswered, with uninformative probabilities. It is never dropped.