A serving platform can have idle GPUs and still fail to absorb a short burst. The requested model is not yet present on those devices, and the waiting requests cannot use their compute capacity. By the time a complete model copy arrives, the queue may already have exceeded the latency objective.
BlitzScale changes when a new instance becomes useful[1]. The Shanghai Jiao Tong University and Huawei Cloud collaboration combines faster weight distribution with cooperative execution between existing and partially loaded instances. Instead of treating model startup as a single unavailable interval, it progressively transfers part of the computation to the new resources.
The relevant objective is therefore not simply a shorter file load. It is earlier additional throughput for requests that would otherwise wait. That distinction explains both the system’s design and the conditions under which its reported improvements should be expected.
Why spare capacity is not ready capacity
An instance must establish an execution environment and populate GPU memory with model parameters. A platform can prepare CUDA contexts and kernels in advance, but tens or hundreds of gigabytes of weights still need to reach the target. Copying from storage may take longer than a short burst can tolerate.
Host-memory caching accelerates this transfer when the destination host already contains the model. However, a platform serving many models cannot assume every host has every weight set. A model that was useful earlier may have been evicted, and scaling across more hosts increases the opportunity for local cache misses.
This creates a mismatch between aggregate resources and accessible resources. There may be enough memory somewhere in the cluster, or a running GPU may already hold exactly the required weights, while the selected destination still takes a slow storage path. BlitzScale makes those existing copies usable as sources for scaling.
What constant host caching actually means
The paper’s O(1) claim concerns the number of host-cache copies required per model as the cluster grows. It does not mean that all model weights occupy constant bytes regardless of the number or size of models. It also does not eliminate the GPU copies needed by active serving instances.
A global parameter manager tracks copies in host memory and on deployed GPUs. A running instance can supply parameters directly. If no GPU serves the model, a host-memory copy can seed the transfer. This removes the requirement to maintain a local host copy beside every possible target.
The design assumes that aggregate memory can hold the required parameter pool and that tracked sources remain available. A failed host may remove the only cached source for a model, requiring recovery and redistribution. Reducing redundant cache copies is a capacity advantage, but it makes source placement and recovery policy more consequential.
Distributing weights without serializing every recipient
Sending a complete model independently from one source to every target can overload that source. BlitzScale instead uses forwarding chains. A recipient forwards early chunks while later chunks are still arriving, allowing multiple links to carry different portions of the model concurrently.
For large transfers on suitable links, this pipeline avoids multiplying the source’s serialization time by the recipient count. It does not make propagation delay, startup, or the slowest link disappear. The useful approximation depends on bulk-transfer behavior and adequate overlap between receiving and forwarding.
The planner groups GPUs connected by fast scale-up links and can distribute different parameter shards across inter-host paths before gathering them locally. It also prefers sources and targets that avoid slower inter-leaf communication when possible. A greedy online plan is used because waiting for a globally optimal plan would itself delay the scale-out response.
The physical network consequently matters to the software policy. A group with fast local GPU communication offers different transfer opportunities from GPUs connected only through PCIe. The paper’s actual evaluation includes both kinds of cluster, rather than assuming every installation has the same intra-host fabric.

Weight traffic competes with the service it is helping
In disaggregated serving, prefill produces KV state that must reach decode instances. Copying model weights over the same outgoing path can delay that existing traffic. A faster model load is not a successful optimization if it increases token latency for requests already in service.
BlitzScale uses knowledge of serving traffic direction when choosing sources and transfer chains. For example, obtaining weights from a decode instance can avoid competing with a prefill instance’s KV transfers. Multiple chains can also reduce cases where a target must both forward weights and send newly produced KV data through the same constrained direction.
This is a topology-aware scheduling opportunity, not a universal claim that full-duplex links eliminate interference. PCIe, memory bandwidth, switches, and other flows can still be shared. The paper simplifies the network with scale-up groups and a leaf-spine model; an operator should validate the actual bottlenecks rather than copy the direction rule without measurements.
Useful execution before a complete copy
Faster copying alone retains a fundamental delay if an instance cannot contribute until all weights arrive. Even overlapping loading with that instance’s own inference does not necessarily solve it: the instance still cannot complete a request until the final required layers become available.
BlitzScale instead pairs the new instance with an overloaded existing instance. The new device executes the prefix of layers it has loaded, then transfers the activation to the existing instance, which executes the remainder. The complete model computation still occurs. No response is produced from an incomplete model.
The old instance now performs less work per request, so its queue can drain faster before the new instance becomes independently operational. As additional layers arrive, the split can move. Eventually both instances hold the complete model and requests can be distributed normally.
This temporary arrangement is distinct from permanently partitioning a model. Its purpose is to use partial readiness during a transition. Whether it helps depends on the work removed from the overloaded instance exceeding the activation-transfer and coordination costs introduced by cooperation.
Why the currently loaded prefix is not always the best split
A straightforward scheduler would give the new instance as many currently available layers as possible and immediately send every request onward. Early in loading, however, that prefix is short. Many requests can still accumulate at the old instance because nearly all computation remains there.
ZigZag scheduling considers the fact that more layers are arriving while requests wait. If the old instance is busy, a request can remain available for additional work on the new instance after the next layer loads. The scheduler revisits such requests, moving more computation away from the congested peer.
The paper develops a formulation for the pipeline split and an implementation that avoids solving the optimization problem for every decision. A priority queue combines request ordering with the availability of the next layer. The existing instance takes eligible work when it can proceed, while the new instance continues processing what has become executable.
The mechanism should not be interpreted as arbitrary intentional delay improving every workload. Its benefit comes from replacing waiting behind an overloaded peer with useful work on a device whose capability is increasing. When there is little queueing or transfers dominate, additional coordination can have less value.
Decode scaling requires a different transition
Directly live-scaling a decode instance can create incoming-bandwidth contention between parameter loading and KV transfers. The paper handles this by changing some existing prefill instances into decode instances and concurrently replenishing prefill capacity through live scaling.
This works because the two roles use the same model parameters, even though their execution and state requirements differ. It also shows why autoscaling cannot be reduced to counting identical empty slots. The role assigned to a ready instance affects which traffic paths and transition costs the system can use.
The implementation monitors token processing and KV occupancy to trigger scaling, but the authors separate this mechanism from the broader policy of when and how far to scale. Workload-sensitive policies and changes within a model’s parallel configuration remain areas for further study, including more detailed exploration of MoE behavior.
What the test clusters establish
The evaluation uses a 32-GPU A800 cluster with NVLink and a 16-GPU A100 cluster with PCIe-based local connectivity. Both have 100 Gbps inter-host RDMA in the reported testbed. These are the experimental conditions, distinct from faster illustrative network rates discussed elsewhere in the paper.
BurstGPT, AzureCode, and AzureConv traces are rescaled while preserving their temporal pattern, with average load set relative to cluster capacity. The main comparison pairs particular models and traces rather than testing every possible combination. Real trace origins make the bursts meaningful, but this remains a controlled replay study rather than a demonstrated fleet deployment.
ServerlessLLM provides the autoscaling baseline. An additional AllCache variant assumes every needed host-memory copy is available, separating local cache misses from other loading limitations. The same scaling policy is applied to these variants, helping isolate the effect of the scaling mechanism under that policy.
Latency reductions that should not become one headline
For the reported BurstGPT and 72B-model case, BlitzScale reduces P95 first-token latency by 75.5% relative to ServerlessLLM. The corresponding token-spacing reduction is 7.4%. Against AllCache, the reductions are smaller: 21.1% for first-token latency and 5.1% for token spacing.
The approximately 94% token-spacing improvement comes from the AzureCode 8B case, where the timing of bursts and cache eviction make slow scaling particularly costly. It is not the expected decode improvement for every model. In other evaluated workloads, token-spacing gains are much smaller because decode capacity can be prepared earlier.
A separate 24B experiment scales six prefill instances in about 1.2 seconds, versus roughly 2 seconds for AllCache, and shows useful contribution before loading completes. This illustrates the distinction between first useful capacity and complete readiness, without establishing a universal subsecond model-start guarantee.

GPU time is not the same as purchased capacity
In the AzureConv 24B comparison, the plotted GPU-time total is 51.62% of the fully provisioned baseline for BlitzScale, versus 71.08% for ServerlessLLM. GPU time is the integral of allocated GPUs over the replay interval. The difference between those plotted percentages is a percentage-point difference, not the same numerical relative reduction.
Using less GPU time while meeting the experiment’s latency objective supports a resource-efficiency argument. It does not directly prove the same reduction in fleet hardware, electricity, or total operating expense. Those depend on whether released GPUs can serve other models, whether their demand peaks coincide, and what reserve capacity is required for failures.
Likewise, the fully provisioned system can still provide the best raw latency. Autoscaling seeks an acceptable latency objective with fewer resources, not necessarily the minimum latency regardless of cost. The objective and tolerated violations must remain explicit when comparing these operating points.
Conditions for an operationally useful transition
A deployment should measure at least three timelines separately: detecting the need to scale, producing the first useful extra capacity, and reaching full independent readiness. Faster parameter movement cannot repair a trigger that arrives after the queue has already become unacceptable.
Source failure, model-version consistency, and interrupted transitions also need tests. A request must never continue through layers from a different weight version simply because both are available. The global parameter pool and temporary layer split make identity and transition state part of correctness, not only performance bookkeeping.
Finally, test weight distribution while the network carries representative KV and other serving traffic. The useful bandwidth is what remains on the actual path during the burst, not an unloaded link’s advertised rate. BlitzScale’s main insight is that partially prepared resources can be productive; its practical value depends on making that productivity arrive without destabilizing the work already running.
Sources and rights
This independently written analysis reviews the OSDI 2025 paper. Reported measurements come from the paper; deployment implications are our interpretation. Original paper copyright remains with its authors, © 2025. Figures were created for this article. The conceptual hardware material is not a photograph or wiring specification of the evaluated systems.