JevServe-bench

Archived board: v0.1 as published on 30 Sep 2026 · the current board

Leaderboard

Decision models, measured inside the serving stack

A decision model reads a request and answers typed questions with probabilities; it writes no text. JevServe-bench puts one in charge of eight decisions an LLM serving system makes, from routing and scheduling to guardrails, and scores what each component then achieves. On every task, 0 is what the system does without the model and 100 is an oracle that knows the outcome in advance.

01Leaderboard

    The average is the plain mean of the eight task scores. Select a column to sort by it, or a model's name for the raw numbers behind its scores.

    p50 is the decision model's median single-request latency on the submitter's setup, so it is reported, not compared. It enters one score: in the scheduling task every request waits that long for its prediction.

    02Eight decisions on the request path

    Each task is one decision a serving system makes, with fixed questions and a fixed downstream rule, so the decision model is the only thing that changes between entries. In request order:

      03How to read a score

      A share of the achievable gain

      A score is how much of the gap between the serving default and the oracle the component closes. Below 0, it does worse than the default.

      No thresholds

      Classification tasks use AUROC, and the thinking and cascade tasks act on requests in the order the model ranks them, so a model is judged on its ordering rather than on a tuned cutoff.

      The same score on any hardware

      Tier A replays logged outcomes and human labels, so the same answers score the same on any machine. Tier B runs a vLLM scheduler simulator calibrated on an H200 serving Qwen3-8B, within 2–4% of the real server. The measured decision latency is the only hardware-dependent input.

      Failures count

      A request that still fails after retries is scored as unanswered, with uninformative probabilities. It is never dropped.