An interactive coding assistant may need tens of tokens within hundreds of milliseconds. A data-wrangling task can accept slower output if it receives more total capacity. Continuous batching places both requests in the same iteration and gives them nearly the same per-token cadence. Priority can reorder work, but shrinking the whole batch for one strict request sacrifices throughput for everyone.
AdaServe uses speculative decoding as a per-request rate control. A draft model proposes several future tokens; the target model verifies them without changing the output distribution. Instead of maximizing accepted tokens uniformly, AdaServe chooses a speculation tree that provides enough speed for each request’s service-level objective while conserving GPU work for the batch. Different deadlines receive different tree shapes.
Uniform iterations are the wrong control unit
Autoregressive decoding normally produces one token per request per target-model iteration. Continuous batching keeps the GPU occupied by adding and removing requests, but every active sequence advances at the batch’s iteration time. A strict request cannot run faster unless the system reduces the batch or gives it separate resources. Both choices strand capacity when relaxed requests could coexist.
Speculative decoding can advance several tokens in one target-model verification. The number and branching of draft candidates determine expected progress and cost. A large tree may accelerate a request with a high acceptance rate, but it consumes verification slots that could serve others. The optimal tree therefore depends on request SLO, draft quality, sequence state, batch composition, and GPU throughput.
AdaServe formulates this selection as a constrained optimization. A hardware-aware model estimates how many candidate tokens the target GPU can verify in parallel. The scheduler assigns a tree to each request so its expected token rate meets the objective while total goodput is maximized.

Separating speculation from selection
Running the draft model independently for every candidate branch can erase the benefit. AdaServe separates three stages. Speculation generates candidate material, selection chooses the tree that fits current objectives and hardware capacity, and verification evaluates selected candidates with the target model. Pipelining these stages lets draft work overlap and avoids rebuilding every request’s plan serially.
Selection is the policy point. A request near its deadline can receive a tree expected to produce more accepted tokens. A relaxed request can use a smaller tree or ordinary decoding, leaving verification capacity available. The scheduler does not change token correctness; rejected draft paths are discarded and target-model sampling preserves the original distribution.
The system adapts tree parameters as request mix and load change. Static profiles describe GPU processing capacity, while runtime observations reveal acceptance and queue conditions. Adaptation matters because the same tree that helps at low load can overload verification at high load. The control loop must change quickly enough that an SLO regime does not end before the policy converges.
Results across several SLO views
AdaServe reduces SLO violation rate by up to 4.3x and improves goodput by up to 1.9x over the best-performing baseline in the reported workloads. As the fraction of strict requests rises, SLO satisfaction reaches 1.5x the baseline while goodput rises by 64%. Under strict time-per-output-token objectives, the gain in goodput reaches 1.38x.
These are different experiments, not additive multipliers. Violation reduction depends strongly on a baseline’s initial miss rate. Goodput counts output that satisfies the selected SLO and is therefore more useful than raw tokens per second, but its value still depends on the exact objective. A request can satisfy time per output token while missing time to first token, or the reverse.
The study uses diverse service traces and model configurations, but a deployment must reproduce its own acceptance rates and arrival distribution. Speculative decoding gains shrink when the draft model predicts poorly, the target is small, or draft compute competes for the same bottleneck. They can grow when verification efficiently processes many candidates and latency classes differ materially.
The cost hidden by accepted-token ratios
Acceptance rate alone does not measure efficiency. A wide tree can contain many rejected tokens, consuming draft and verification compute for little progress. The relevant quantity is SLO-valid target tokens per unit of GPU time, including draft execution, selection, KV-cache access, and discarded branches.
Memory also changes. Candidate branches need temporary token and KV state. A tree selected for a strict request can reduce batch capacity or increase memory fragmentation. The hardware-aware model should include memory limits, not only arithmetic throughput. If the system handles an out-of-memory event by shrinking all trees after a miss, tail latency can worsen precisely during bursts.
Power is another constraint. Speculation intentionally computes work that may be rejected. It can reduce completion time while raising joules per accepted token. Operators with rack power limits should include energy or instantaneous power in the optimization rather than assume unused tensor-core capacity is free.
Fairness among latency classes
Customizing a tree per SLO makes priority explicit. A strict class can consume more speculative resources than relaxed work. If admission exceeds capacity, no tree assignment can satisfy everyone. The system needs a policy for rejection, degradation, and minimum progress rather than hiding overload behind average goodput.
SLOs should come from service value, not user-selected labels without cost. Otherwise every tenant requests the strictest class. Pricing, quotas, or workload identity must connect the objective to an allocation right. The scheduler should report achieved latency and consumed verification capacity by class so operators can audit that policy.
Starvation can occur when strict traffic remains high. Relaxed requests may meet a loose per-token limit yet make little total progress. Maximum queue age or completion-time objectives can supplement token SLOs. A multi-SLO system is complete only when it defines what happens after demand exceeds the feasible frontier.
Comparison with priority and disaggregation
Priority scheduling can run urgent requests first but does not change the one-token-per-iteration structure. Preemption and smaller batches lower latency at a throughput cost. AdaServe changes per-request progress within a shared verification step, creating another control dimension.
Disaggregated prefill and decode can scale phases independently, but decode workers still face mixed token objectives. AdaServe can operate inside that decode pool. Its speculation pipeline may use separate draft resources, although network and KV transfer then enter the latency budget. Colocation avoids transfer but shares compute and power.
Model routing provides another alternative: send strict requests to a smaller model. That changes output quality, while lossless speculative decoding preserves target-model outputs. A service can combine both, but it must distinguish quality tradeoffs from systems acceleration.
A production acceptance test
The test matrix should cross SLO class, prompt and output length, model, draft model, request rate, and burst shape. Report time to first token, time per output token, completion time, violation rate, raw throughput, goodput, rejected candidate work, memory, and power. Per-class p95 and p99 results matter more than one aggregate mean.
The adaptation test should change the request mix abruptly and measure convergence. It should include draft-model slowdown, acceptance-rate drift, and memory pressure. A stable-trace optimum is insufficient if the control loop misses deadlines during every transition.
Correctness must also be tested. Candidate rejection and sampling should match the target distribution within the algorithm’s guarantee. Logging needs to expose tree choices without retaining sensitive prompts. Fallback to ordinary decoding should be safe and observable when profiles are absent or the optimizer cannot find a feasible plan.
The system decision
AdaServe turns speculative decoding from a uniform acceleration technique into an allocation mechanism. The tree attached to a request represents how much extra compute the service is willing to spend to meet that request’s token deadline. A shared verification pipeline then converts that policy into GPU work.
The design is compelling when one model serves genuinely different latency classes and speculation has enough acceptance to move requests several tokens at a time. It is less useful when SLOs are uniform, draft quality is low, or memory and power already limit capacity. The purchase and deployment metric is therefore goodput by class under changing load, with rejected work visible. A 1.9x peak gain matters only if the strict and relaxed users represented in that denominator match the real service.
Source and copyright note
This article is an independent editorial digest of the EuroSys 2026 paper[1]. The prose and figure were created anew for Silicon & Systems; no paper figure or table was reproduced. Results retain the authors’ workloads, SLO definitions, models, and baselines. Copyright in the original paper is held by its authors and publication rights are licensed to ACM (2026).