JevServe-bench

Leaderboard

JevServe-bench

JevServe-bench evaluates System One decision models across LLM system decisions, including routing, caching, scheduling, verification, guardrails and tool use.

Requestsa prompt, a response or an agent's step
Decision modeltyped questions in, probabilities out; no text

    Leaderboard

      Tasks

      The decisions

      Each task is one decision a serving system makes, with fixed questions and a fixed downstream rule, so the decision model is the only thing that changes between entries. On every task, 0 is what the system does without the model and 100 is an oracle that knows the outcome in advance.

      Methodology

      How the scores work

      Every task scores what a serving component achieves when a decision model makes its decision, on a scale from the system's default (0) to an oracle that knows the outcome in advance (100).

      What 0 and 100 mean

      A share of the achievable gain

      A score is how much of the gap between the serving default and the oracle the component closes. Below 0, it does worse than the default.

      No thresholds

      Classification tasks use AUROC, and the thinking and cascade tasks act on requests in the order the model ranks them, so a model is judged on its ordering rather than on a tuned cutoff.

      Averages and categories

      The average is the plain mean of the task scores. A category's score is the mean of its tasks; the average is not a mean of the categories.

      Every score has an interval

      A 95% interval: a bootstrap over the task's items, a prompt's responses resampled together, or for routing and scheduling the spread over data splits and arrival draws. Ranks and bold scores compare models on the same resamples, paired, and call two models tied unless one is ahead beyond the 95% interval of their difference; tied models share a rank.

      Open problems

      Decision latency

      p50 is the decision model's median single-request latency on the submitter's setup, measured after a warm-up, so it is reported rather than compared. It enters one score: in the scheduling task every request waits that long for its prediction.

      Failures count

      A request that still fails after retries is scored as unanswered, with uninformative probabilities. It is never dropped.

      Add your model

      1. Serve it behind POST /v1/systemone

        Any HTTP server that speaks TypeSafe's System One API works, which hosted Jev and the Jev-compatible open models already do. Each request carries the text and the questions; the reply carries probabilities.

        # an example exchange
        POST /v1/systemone
        {"state": "Prove that there are infinitely many primes.",
         "questions": {"needs_thinking": {"type": "noul", "instructions": "Does answering this request
           correctly require thinking step by step at length before responding, rather than answering directly?"}}}
        
        # the reply: P(yes) for a yes/no question
        {"answers": {"needs_thinking": {"noul": 0.91}}}
      2. Run the benchmark

        Prepare the data first (README, “Data”). A full run asks every question once and caches the responses, so an interrupted run resumes; --limit 50 checks the whole pipeline on 50 items per task.

        git clone  && cd jevserve-bench
        uv venv .venv && uv pip install --python .venv/bin/python -e ".[prepare]"
        .venv/bin/jevserve run --backend http --base-url http://127.0.0.1:8000 --model my-model --name my-model
      3. Send the results

        Open a pull request with results/my-model/: scores.json, environment.json and the answer files. We re-score the answers and add a model card before listing it.