Experiment report

Queue ordering by predicted output length on a real vLLM server

Does serving the shortest predicted request first help a real vLLM server, and how close does Jev get to knowing the true length?

JevServe-Bench · Blog draft · Leaderboard

Summary

SLO attainment and P90 time to first token against arrival rate for FCFS, shortest prompt first, Jev and the true length.
Figure 1. SLO attainment and P90 time to first token against arrival rate, arrival draw 0. Hollow markers: the backlog grows. LTR is in the tables below.

Setup

ItemValue
Hardware8x NVIDIA H200 (143,771 MiB) in one host, 192 CPU cores; one server per GPU, up to 8 runs at once
EnginevLLM 0.30.0 (V1), OpenAI-compatible server
ModelQwen/Qwen3-8B, revision b968826d9c46dd6066d109eabc6255188de91218, bf16
Server flags--max-model-len 24576 --no-enable-prefix-caching --num-gpu-blocks-override 5250 --scheduling-policy fcfs or priority
Effective configurationKV cache 84,000 tokens (5,250 blocks of 16); chunked prefill; 8,192 batched tokens per step; 1,024 sequences (vLLM 0.30 default on H200); recompute preemption; CUDA graphs up to batch 512; no reasoning parser
Workloadfirst 2,000 test-split ShareGPT first-turn prompts (single-turn chat)
Output lengthsQwen3-8B's own thinking-on generations logged earlier (temperature 0.6, top_p 0.95, top_k 20, max_tokens 16,384; thinking tokens counted), forced by max_tokens = logged length and ignore_eos: true (temperature 0 at replay)
Length statisticsprompt mean 245 / p50 44 / p99 2,862 / max 4,059 tokens; output mean 2,402 / p50 1,692 / p90 4,932 / p99 16,384 (cap; 21 requests)
ArrivalsPoisson: one unit-rate trace per seed (2,000 exponential gaps rescaled to exactly 1 request/s), divided by the rate; identical arrival times for every policy at a given seed and rate
Clientone process per server (httpx async, streaming); dispatch lag 0.6-0.7 ms median, under 75 ms max
Warm-up8 training-split prompts x 16 tokens before each run

Policies

All but FCFS use --scheduling-policy priority with per-request priority = round(1000 x log2(predicted output length)), smaller first. Waiting requests never preempt running ones; under memory pressure vLLM preempts the lowest-priority running request.

PolicyPredicted lengthDecision latency (added before sending)
FCFS (vLLM default)none0
Shortest prompt first (spf)prompt token count0
LTR ranker (ltr)LTR-style ranker reimplemented after Fu et al. (NeurIPS'24): OPT-125M, ListMLE, trained on 13,914 training prompts' Qwen3-8B thinking-on lengths20 ms (assumed)
Jev 1.13 (jev)Jev's answers to 6 length questions, mapped to log2 length by ridge regression fit on the same 13,914 training prompts92 ms (measured median)
True length (oracle)the logged output length0

Prediction quality on the 6,086 thinking-on test prompts (Kendall tau): Jev 0.559, LTR 0.545, prompt length 0.114. A second answer sampled for the same prompt reaches 0.725, the practical ceiling of any predictor that sees only the prompt.

Metrics

Runs

RoundRunsConfigurationsStatus
08KV 84k: FCFS / Jev / true length at 1.225 and 1.472; KV 190k: FCFS / Jev at 2.099; draw 0superseded (one client process for 8 servers lagged; prefix caching on); used to find the simulator's biases
calibration6decode and prefill step times, TTFT probesused to fit the simulator's H200 profile
124FCFS 1.05 / 1.15 / 1.25; Jev, LTR, true length 1.30 / 1.42 / 1.54; draws 0 and 1done
28Jev, LTR, true length 1.36; FCFS 1.10; draws 0 and 1done
315FCFS 1.20 / 1.30 / 1.42 / 1.54; shortest prompt first 1.10-1.54; Jev, LTR, true length 1.10 / 1.20; draw 0done
3, draw 115the same for draw 1deferred (11 placeholders, 4 stopped)

Each run takes 25-33 minutes. Rounds 1-3 total 47 completed runs.

Results: every policy at the same rates (arrival draw 0)

SLO attainment (* = the backlog grows)

Policy1.101.201.301.421.54
FCFS (vLLM default)92.0%87.8%75.4%41.5%*36.1%*
Shortest prompt first99.3%98.9%96.5%92.5%*88.4%*
LTR ranker99.5%99.2%98.1%96.6%*95.5%*
Jev 1.1399.4%99.2%98.6%97.2%*96.0%*
True length (oracle)99.9%99.9%99.5%99.2%*98.7%*

Mean TTFT (s) (* = the backlog grows)

Policy1.101.201.301.421.54
FCFS (vLLM default)0.521.052.3116.11*37.91*
Shortest prompt first0.170.261.5612.10*33.22*
LTR ranker0.090.201.334.57*6.45*
Jev 1.130.150.380.682.65*3.89*
True length (oracle)0.050.050.430.96*1.10*

P90 TTFT (s) (* = the backlog grows)

Policy1.101.201.301.421.54
FCFS (vLLM default)0.102.808.2743.02*101.27*
Shortest prompt first0.050.050.070.13*6.14*
LTR ranker0.070.070.110.13*0.17*
Jev 1.130.140.150.170.18*0.23*
True length (oracle)0.050.050.060.06*0.09*

P99 TTFT (s) (* = the backlog grows)

Policy1.101.201.301.421.54
FCFS (vLLM default)10.519.228.593.8*163.6*
Shortest prompt first0.42.752.9445.2*764.5*
LTR ranker0.40.622.093.3*224.1*
Jev 1.130.50.62.739.4*171.6*
True length (oracle)0.30.30.40.5*2.1*

Mean end-to-end latency (s) (* = the backlog grows)

Policy1.101.201.301.421.54
FCFS (vLLM default)19.020.224.144.3*69.5*
Shortest prompt first18.419.123.338.3*65.8*
LTR ranker18.519.526.032.0*34.6*
Jev 1.1318.619.624.428.3*33.9*
True length (oracle)17.919.422.424.4*26.0*

Results: goodput

PolicyDraw 0, realDraw 0, simulatedDraw 1, realDraw 1, simulated
FCFS1.127 (interpolated)1.1591.219 (interpolated)1.217
Shortest prompt first1.30-1.42 (backlog)-not run-
LTR ranker1.30-1.36 (backlog)1.3601.42-1.54 (backlog)1.490
Jev 1.131.30-1.36 (backlog)1.4091.42-1.54 (backlog)1.510
True length1.36-1.42 (backlog)1.419>= 1.541.575

The calibrated simulator's goodput averaged over both draws: FCFS 1.188, shortest prompt first 1.394, LTR 1.425, Jev 1.460, true length 1.497 req/s (+17%, +20%, +23%, +26% over FCFS).

Results: real server vs calibrated simulator

Real and simulated latency curves for FCFS, the LTR ranker, Jev and the true length.
Figure 2. Rounds 1 and 2: the real server (solid, mean of two arrival draws) against the calibrated simulator (dashed).

Interpretation

  1. Admission packing: in this memory-bound setting, the largest part of the gain over FCFS comes from not blocking admission behind a request that does not fit. Shortest prompt first, which knows nothing about output length, already moves attainment at 1.42 req/s from 41.5% (FCFS) to 92.5%.
  2. Output-length information adds attainment at high load and shortens the TTFT tail. The ordering is clear: true length > Jev ~ LTR > shortest prompt first > FCFS. Jev and LTR are close; Jev has the shorter tail at 1.30-1.42 req/s, LTR the lower P90 at low load because of Jev's decision latency.
  3. Starvation binds before the SLO for every shortest-first policy, from about 1.36-1.42 req/s. That is why their goodput is bracketed and nearly tied. Starvation control (aging) would be needed to push goodput further and to separate predictors by goodput.
  4. The setting matters: these gains need a small KV pool (tens of requests at once). On default deployments with hundreds of requests in the cache, the simulator finds 0-3% goodput gains even with the true length (the simulated deployment sweep).

Caveats

Every run

Show all 47 runs
policyrateseedattainment real (sim)backlog grows real (sim)mean TTFT s: client / server / simP90 TTFT s real (sim)P99 TTFT s real (sim)mean e2e s real (sim)TPOT p50 ms real (sim)preemptions real (sim)
fcfs1.05095.7% (97.2%)no (no)0.22 / 0.21 / 0.120.06 (0.04)5.0 (4.2)18.2 (17.3)7.2 (7.0)408 (296)
fcfs1.05197.7% (98.2%)no (no)0.11 / 0.11 / 0.080.05 (0.03)3.1 (2.3)17.2 (17.1)7.0 (7.0)199 (197)
fcfs1.10092.0% (93.5%)no (no)0.52 / 0.52 / 0.360.10 (0.05)10.5 (9.8)19.0 (18.1)7.4 (7.1)641 (476)
fcfs1.10196.3% (97.0%)no (no)0.23 / 0.22 / 0.190.06 (0.04)8.2 (7.5)18.1 (17.6)7.3 (7.1)341 (297)
fcfs1.15088.2% (90.3%)no (no)0.96 / 0.95 / 0.681.99 (0.76)17.6 (13.9)20.3 (19.1)7.5 (7.2)937 (820)
fcfs1.15195.2% (94.8%)no (no)0.36 / 0.36 / 0.380.07 (0.06)12.8 (12.4)18.1 (18.3)7.2 (7.2)445 (539)
fcfs1.20087.8% (87.1%)no (no)1.05 / 1.05 / 0.992.80 (2.54)19.2 (17.9)20.2 (20.3)7.4 (7.4)1050 (1141)
fcfs1.25083.2% (83.1%)no (no)1.46 / 1.45 / 1.474.93 (4.83)19.3 (19.8)21.7 (21.8)7.5 (7.6)1510 (1577)
fcfs1.25187.7% (88.2%)no (no)0.96 / 0.95 / 0.921.86 (1.57)19.1 (19.4)20.5 (20.4)7.5 (7.6)1284 (1158)
fcfs1.30075.4% (76.4%)no (no)2.31 / 2.30 / 2.258.27 (8.09)28.5 (29.1)24.1 (23.6)7.8 (7.8)2165 (1881)
fcfs1.42041.5% (44.1%)yes (yes)16.11 / 16.11 / 13.5443.02 (38.98)93.8 (79.4)44.3 (41.5)10.1 (10.1)4911 (4729)
fcfs1.54036.1% (37.8%)yes (yes)37.91 / 37.90 / 36.36101.27 (99.27)163.6 (152.1)69.5 (67.0)11.2 (11.1)5531 (5373)
spf1.10099.3% (99.2%)no (no)0.17 / 0.16 / 0.130.05 (0.03)0.4 (0.5)18.4 (18.0)7.3 (7.1)146 (170)
spf1.20098.9% (98.9%)no (no)0.26 / 0.26 / 0.260.05 (0.03)2.7 (3.6)19.1 (18.8)7.5 (7.5)243 (243)
spf1.30096.5% (97.4%)no (no)1.56 / 1.56 / 0.700.07 (0.04)52.9 (24.9)23.3 (21.2)8.5 (8.0)755 (649)
spf1.42092.5% (93.0%)yes (yes)12.10 / 12.09 / 8.400.13 (0.10)445.2 (252.5)38.3 (34.1)8.8 (8.6)1096 (964)
spf1.54088.4% (90.6%)yes (yes)33.22 / 33.22 / 25.576.14 (0.29)764.5 (665.2)65.8 (53.4)9.0 (8.7)1331 (1072)
ltr1.10099.5% (99.7%)no (no)0.09 / 0.06 / 0.060.07 (0.05)0.4 (0.3)18.5 (17.9)7.3 (7.1)181 (144)
ltr1.20099.2% (99.2%)no (no)0.20 / 0.17 / 0.170.07 (0.05)0.6 (0.5)19.5 (19.2)7.5 (7.5)296 (274)
ltr1.30098.1% (98.2%)no (no)1.33 / 1.30 / 0.520.11 (0.07)22.0 (5.6)26.0 (21.6)8.9 (8.2)715 (628)
ltr1.30198.1% (99.0%)no (no)1.11 / 1.09 / 0.410.09 (0.06)7.1 (1.0)23.2 (20.8)8.5 (8.1)601 (419)
ltr1.36097.0% (97.8%)yes (no)1.72 / 1.69 / 1.310.12 (0.09)48.8 (26.4)29.1 (26.2)9.0 (8.7)829 (804)
ltr1.36197.9% (98.2%)no (no)1.78 / 1.75 / 0.720.10 (0.08)47.4 (11.6)26.8 (23.7)9.1 (8.5)833 (681)
ltr1.42096.6% (97.5%)yes (yes)4.57 / 4.54 / 1.890.13 (0.09)93.3 (55.2)32.0 (28.6)9.1 (8.7)806 (802)
ltr1.42197.5% (97.5%)no (no)2.00 / 1.97 / 1.630.12 (0.08)53.7 (32.6)27.3 (25.8)8.8 (8.7)843 (882)
ltr1.54095.5% (96.2%)yes (yes)6.45 / 6.42 / 5.220.17 (0.11)224.1 (199.7)34.6 (33.0)9.0 (8.8)1126 (969)
ltr1.54195.3% (96.2%)yes (yes)7.58 / 7.56 / 5.310.19 (0.14)321.6 (182.0)36.1 (32.5)9.2 (8.8)1287 (1033)
jev1.10099.4% (99.6%)no (no)0.15 / 0.05 / 0.130.14 (0.13)0.5 (0.5)18.6 (18.0)7.3 (7.1)201 (181)
jev1.20099.2% (99.3%)no (no)0.38 / 0.28 / 0.300.15 (0.13)0.6 (0.4)19.6 (19.1)7.6 (7.5)245 (224)
jev1.30098.6% (99.4%)no (no)0.68 / 0.58 / 0.410.17 (0.14)2.7 (0.6)24.4 (21.0)8.9 (8.1)629 (465)
jev1.30199.1% (98.9%)no (no)0.54 / 0.44 / 0.470.16 (0.13)0.9 (1.2)22.1 (21.0)8.4 (8.0)471 (396)
jev1.36097.5% (98.6%)yes (no)1.39 / 1.29 / 0.540.18 (0.15)19.2 (2.2)26.4 (23.2)9.1 (8.5)790 (610)
jev1.36198.0% (98.6%)no (no)0.92 / 0.82 / 0.550.17 (0.14)9.9 (2.9)25.4 (22.3)9.2 (8.4)754 (569)
jev1.42097.2% (98.0%)yes (yes)2.65 / 2.55 / 1.280.18 (0.16)39.4 (16.3)28.3 (25.2)9.2 (8.7)849 (767)
jev1.42197.6% (98.2%)no (no)1.51 / 1.41 / 0.710.17 (0.15)19.0 (11.5)25.7 (24.1)8.9 (8.7)735 (706)
jev1.54096.0% (96.6%)yes (yes)3.89 / 3.79 / 2.850.23 (0.18)171.6 (88.0)33.9 (28.8)9.2 (8.8)1228 (942)
jev1.54196.1% (96.3%)yes (yes)4.08 / 3.98 / 3.990.26 (0.20)130.5 (128.2)34.3 (29.9)9.3 (8.9)1442 (1118)
oracle1.10099.9% (99.9%)no (no)0.05 / 0.05 / 0.030.05 (0.03)0.3 (0.3)17.9 (17.8)7.2 (7.1)82 (86)
oracle1.20099.9% (99.8%)no (no)0.05 / 0.04 / 0.030.05 (0.03)0.3 (0.3)19.4 (18.5)7.7 (7.5)186 (150)
oracle1.30099.5% (99.5%)no (no)0.43 / 0.43 / 0.110.06 (0.05)0.4 (0.5)22.4 (20.4)8.7 (8.1)386 (364)
oracle1.30199.7% (99.6%)no (no)0.35 / 0.34 / 0.070.06 (0.04)0.5 (0.4)21.8 (20.1)8.6 (8.0)356 (313)
oracle1.36099.2% (99.3%)no (no)0.67 / 0.66 / 0.440.07 (0.06)0.5 (0.5)23.4 (21.9)8.8 (8.4)408 (413)
oracle1.36199.4% (99.6%)no (no)0.39 / 0.38 / 0.340.07 (0.05)0.5 (0.4)23.0 (21.0)9.0 (8.3)428 (319)
oracle1.42099.2% (99.2%)yes (no)0.96 / 0.95 / 0.690.06 (0.06)0.5 (0.5)24.4 (23.1)8.8 (8.6)374 (516)
oracle1.42199.4% (99.5%)no (no)0.52 / 0.51 / 0.470.08 (0.05)0.5 (0.4)23.6 (22.6)9.0 (8.6)505 (422)
oracle1.54098.7% (98.9%)yes (yes)1.10 / 1.09 / 1.000.09 (0.07)2.1 (1.2)26.0 (25.1)8.9 (8.7)602 (489)
oracle1.54198.9% (99.2%)no (no)2.20 / 2.20 / 0.990.09 (0.08)0.5 (0.7)28.1 (25.2)9.4 (9.0)663 (653)