An agent can produce each answer quickly and still finish its assignment slowly. It may call a model, inspect a tool result, call the model again, and launch several parallel explorations before combining their outputs. Every return to the inference engine can create another wait. Optimizing a single response does not necessarily optimize the program that needs all those responses.
Agentix changes the scheduling unit without requiring the agent to reveal its future execution graph[1]. The NSDI 2026 paper, coauthored by researchers at UC Berkeley, Google DeepMind, and Shanghai Jiao Tong University, tracks the service already received by a program. Subsequent LLM calls inherit that history rather than always entering the queue as unrelated new work.
This is a serving-system prototype, not evidence that Google runs this scheduler in a production agent fleet. Its useful contribution is a concrete explanation of accumulated queueing delay and a mechanism to reduce it. The evaluation also makes clear that headline throughput depends on which inference optimizations the comparison system already enables.
The queue does not see the user’s task
Request-level scheduling receives a stream of prompts. It can observe arrival time, token counts, cache occupancy, and whether a call has finished. Without a program identifier, it cannot distinguish a genuinely new short assignment from the twentieth call of an already expensive assignment. Those calls may look identical at the API boundary while having different implications for completion time.
A first-come policy can let a long decoding call occupy a batch slot while shorter calls wait. Preemption addresses that problem by periodically reconsidering which calls should run. However, a scheduler that gives every new call the highest priority creates a second problem: a long program repeatedly returns with fresh calls and repeatedly receives newcomer treatment.
Agentix separates these two effects. One arises inside the current call; the other arises across calls belonging to the same program. Improving only the first can leave the user’s short program waiting behind the continuing activity of longer programs. A request scheduler therefore needs more context than the age of its current request if it is expected to optimize completed assignments.
The distinction also affects utilization. Finishing one call can release the next dependent call, creating additional batching opportunities. Conversely, delaying a dependency can leave later work unavailable even when it would have fit efficiently with other requests. The paper’s argument is not simply that latency matters alongside throughput. Reducing the right waiting time can change how much runnable work the engine sees.
A session record as scheduling input
Agentix associates calls with a persistent session and maintains a process table. The table records accumulated execution, waiting, active calls, and engine placement. Programs still run outside the inference service and issue calls as their control flow develops. The serving layer learns their observed behavior incrementally rather than requiring a complete static workflow at submission.
For a sequential program, the relevant history is the execution time consumed by completed LLM calls. A new call begins with that accumulated service value. A program that has already received substantial computation therefore does not reset its standing merely by crossing another API boundary. The current call can still be demoted as it continues executing.
This mechanism does not predict exactly how much work remains. A program with little prior service can later become enormous, and a program with a large history can be one call from completion. The policy uses observed service as an information-limited scheduling signal, not a claim to know remaining processing time. The appendix’s comparison with a simulator that knows future work leaves a performance gap.
Session correctness becomes part of system correctness. Programs need identifiable starts, completions, and errors so records do not persist indefinitely or disappear prematurely. An operator must also decide what constitutes one program: a user conversation, a research request, or a background workflow. Changing that definition changes the history presented to the scheduler and can change the resulting priorities.

Sequential history and parallel progress
The sequential policy, PLAS, ranks work using service accumulated by the originating program. This resembles least-attained-service scheduling, but the accounting survives individual call completion. Calls are not treated as an endless stream of unrelated jobs. That is the specific change needed to prevent repeated priority resets.
Parallel programs complicate the accounting because summing every thread’s service can overstate the sequential work determining completion. Several calls may execute concurrently, after which a join waits for the slowest branch. A useful priority signal must represent progress along dependent execution, rather than equate total work with elapsed completion time.
ATLAS maintains a scalar reflecting the longest observed service path for each program. Active calls inherit the program’s current value, and completed work updates it when the new value is larger. This is a practical approximation based on information already observed. It is not advance knowledge of the eventual critical path, nor a requirement to supply every future dependency edge.
Giving related parallel calls compatible priorities also helps avoid leaving one branch behind while unrelated work advances. However, the benefit depends on the available batching capacity and the behavior of the other programs. A wide fanout does not create more GPU memory, and a program with many active calls still needs safeguards against overwhelming the serving engine.
Preemption needs a memory policy
Agentix discretizes priorities into several queues instead of continuously swapping the next nominally best call into execution. Calls receive time quanta and move between queues as they consume service. This limits excessive context changes that could turn a theoretically attractive priority rule into expensive round-robin behavior.
Long programs also need a way to make progress. The scheduler compares accumulated waiting with accumulated service and promotes work when their ratio crosses a threshold. That threshold expresses a tradeoff between favoring short completion paths and protecting delayed programs. It should not be confused with a complete customer-level fairness or billing policy.
Preempted calls retain KV state that occupies memory. When that state moves between GPU and host memory, fragmented transfers can make frequent switching expensive. Agentix gathers KV blocks for bulk transfer and reduces scheduling frequency through multistep execution. The reported system gains include these implementation choices, not only the priority formula.
This matters when assessing adoption. Reimplementing the queue order while leaving a costly swap path unchanged may produce a different result. Conversely, a newer engine with better swapping, caching, and scheduling can reduce the advantage available to the proposed policy. The right comparison is the complete serving stack under the same memory limit, workload, and latency requirement.
Locality across engine replicas
Calls within a program often reuse accumulated context. Routing a long follow-up prompt to the engine holding its prefix can avoid repeating substantial prefill work. Balancing every call independently may distribute the queue evenly while wasting that reusable state. Session identity gives the router a relatively inexpensive proxy for this locality.
Agentix treats short prompts differently. In the studied traces, shared system prompts account for much of their reusable input, so a short call can go to a less-loaded engine with limited locality loss. Longer calls tend to remain on their program’s assigned engine. The implementation uses a 2,048-token cutoff derived from those traces.
That cutoff is not a property of GPU hardware or a universal threshold for agents. A service with large tool schemas, changing system prompts, retrieval-heavy context, or model switching can have a different reuse pattern. Program identity also does not prove that two prompts actually share a prefix. The routing rule should be tested against observed cache hits and queueing time.
The multi-engine evaluation compares routing policies while holding the program-aware scheduler in place. This separates the benefit of locality-aware routing from the larger benefit of replacing the baseline serving configuration. The paper’s main evaluation reports up to 1.4× throughput against simpler routers and behavior close to a prefix-aware alternative, under its tested replica arrangements.
What the throughput comparison measures
The testbed uses A100 GPUs with 80 GB each. Llama 3.1 8B, Llama 3.1 70B, and Falcon 180B run on one, four, and eight GPUs, respectively. Workloads replay complete ShareGPT conversations, function-calling sequences from BFCL, and tree-search programs from LATS, plus an equal mixture of these program classes. Program arrivals follow a generated Poisson process.
These are materially different workloads. The reported averages are about 6.66 calls per ShareGPT program, 10.75 for BFCL, and 159.7 for LATS. A scheduler sees different levels of repeated queue entry, parallelism, and prefix reuse in each. An equal mixture in an experiment is not a measurement of the traffic mix at a commercial provider.
The principal latency metric divides each program’s response time by its generated tokens and then averages those values. For parallel programs, the completion-path response time is divided by tokens across the threads. This is not ordinary time per output token for an individual streaming response. It also does not make a token-normalized ratio interchangeable with every user’s absolute completion-time requirement.
The baseline is vLLM 0.6.1, with an additional optimized configuration enabling chunked prefill, prefix caching, and multistep scheduling. Agentix’s large gains against the first configuration partly reflect missing optimizations there. On the mixed workload, the paper reports maxima of 15× against basic vLLM and 5× against the optimized version. Those are different comparisons, not two estimates of one universal scheduler speedup.

Tail results and offline results answer additional questions. The paper reports better P95/P99 behavior in seven of eight tested tail scenarios against the compared optimized schedulers. In a separate offline batch experiment, makespan falls by 10–40%. Neither result means every long program becomes faster, and neither should be substituted for an interactive first-token measurement.
Where the service benefit can disappear
Agentix controls the inference service, not the external world. A program waiting on a slow search API, a database transaction, a browser, or a human remains delayed outside the LLM engine. If those intervals dominate the task, a substantial improvement in serving queues can produce a modest improvement in user-visible completion.
The session abstraction also requires operational discipline. Retries must not accidentally create unrelated fresh identities, while independent tasks should not be merged into one indefinitely expensive session. Authentication and tenant accounting need to prevent priority manipulation. These are deployment questions derived from the design, not capabilities demonstrated by the paper’s throughput plots.
A practical pilot should preserve the current engine’s optimizations and add program identifiers to telemetry first. Measure queueing, model execution, KV movement, and external waits separately. Then compare completed programs under fixed output-quality checks, absolute completion deadlines, and streaming latency targets. This reveals whether repeated queue entry is an important cost before replacing a scheduler.
We believe the central lesson is that agent infrastructure needs an accounting unit larger than one prompt. Agentix provides a concrete way to carry past computation into the next scheduling decision without requiring knowledge of the future. Whether that becomes more useful capacity depends on the actual program mix, the memory cost of preemption, and how much of the user’s wait remains inside the serving system.
Sources and rights
This is an independently written editorial analysis of the final NSDI 2026 paper. The reported measurements belong to that evaluation; deployment judgments are our analysis. Copyright in the original paper remains with its authors under the USENIX publication terms, © 2026. All figures here are original explanatory graphics; no paper figure, table layout, or publisher artwork is reproduced.