GPU colocation is usually framed as a compute-scheduling problem, but memory is the resource that makes inference and training difficult to mix. An online model may use little compute between requests while its weights and working set remain resident. A training job can consume spare cycles, yet its optimizer state, activations and temporary buffers fill whatever HBM is left. When an inference burst arrives, reclaiming that memory by swapping or stopping the job can take hundreds of milliseconds or seconds, long enough to violate the first request’s latency objective.

SIRIUS, presented at USENIX ATC 2025, makes a narrower and more practical trade[1]. Inference receives priority over the whole GPU. Training uses memory and compute that inference does not currently need, but it exposes points inside gradient computation where its allocation can shrink within milliseconds. The next batch is resized to fit the smaller budget. This is a memory handover, not a process replacement.

Why fixed partitions waste the quiet periods

MIG and static spatial partitioning isolate resources well[2]. If a service reserves 75% of HBM for inference, however, that capacity remains unavailable to training even when only a few models are active. Reserving 50% increases training throughput but can force an arriving model to load from host memory or can leave too little KV-cache space for an LLM. The correct split moves with request rate, model popularity and sequence length.

Unified-memory oversubscription appears dynamic but moves pages across PCIe under pressure. The ATC paper measures severe tail-latency and SLO loss when inference competes with training through this path. Whole-task switching has another problem: a training batch may run for hundreds of milliseconds, and stopping it safely can require discarding completed work or waiting until the batch ends. Both mechanisms react at a time scale longer than an online service can tolerate.

SIRIUS exploits a property of backpropagation. Gradient computation proceeds through layers and produces intermediate state whose lifetime ends before the full batch completes. Once a selected segment finishes, training can release activation and workspace memory without corrupting parameters. The control plane tells the training runtime how much memory inference needs, and the runtime reaches a safe adjustment point in a few milliseconds rather than waiting for an entire batch.

Four steps in the handover

First, SIRIUS tracks which allocations are live and which gradient segments can complete under the current budget. It does not evict arbitrary CUDA pages. Second, training frees memory at an explicit boundary and discards only the small amount of unfinished work that cannot be preserved. Third, inference allocation and zero-filling bypass slow general-purpose paths where possible, while model loading is pipelined. Finally, the training job chooses a smaller effective batch for subsequent work.

When inference demand recedes, the process reverses. Training increases its batch and consumes the released memory again. Across multiple GPUs, different devices may face different inference loads, so SIRIUS distributes a training batch unevenly instead of forcing every data-parallel worker to use the same local batch size. Gradients still represent the intended global batch, while memory is not stranded on a lightly loaded GPU.

The policy reserves a watermark for likely inference arrivals and keeps recently idle models live for a configurable interval. A large reserve lowers cold starts but depresses training throughput. A long liveness interval has the same trade. SIRIUS uses an SLO-aware model to select these values from observed arrivals rather than treating maximum inference demand as a permanent partition.

SIRIUS’s GPU-memory handover shown as an original editorial figure. Under normal load, elastic training occupies most memory. During an inference burst, the runtime finishes an adjustable gradient segment, releases live allocations safely, reserves and zeroes inference pages, then resizes the next training batch. The published scope covers data-parallel elasticity; pipeline and tensor-parallel reshaping remain future work. Original figure created for this article.

The evaluated workload

The main evaluation uses NVIDIA A100 GPUs and six DNN inference models compiled with TVM, duplicated into as many as 56 model instances. Arrival patterns include light, heavy, bursty and skewed traces derived from the Microsoft Azure Functions workload. Swin Transformer is the colocated training job. Each run lasts 300 seconds, and the inference SLO is four times the model’s standalone execution latency. Metrics include P99 latency, the share of requests meeting the SLO and training samples per second.

Baselines cover task switching, 50% and 75% static memory partitions, unified memory with MPS, and inference alone. This breadth matters because each baseline protects a different resource. A static partition avoids direct interference but wastes reserved capacity; switching reallocates capacity slowly; unified memory maintains allocations but pays PCIe movement; inference-only establishes the upper bound on latency compliance without useful background work.

Across the tested single-GPU traces, SIRIUS improves SLO compliance by an average 57.0% and up to 97.0% over the compared colocation methods. Training throughput rises by an average 2.2× and up to 13.7×. Relative to inference-only, it retains 95.3% average SLO compliance and reaches as high as 98%. These results do not say that training is free. They say that using the idle intervals costs a small portion of the protected service level under the evaluated arrival process.

On four GPUs, SIRIUS reports 43% higher SLO compliance and 6.1× higher training throughput on average than the compared colocation methods. The batch-distribution mechanism is especially important when inference load is asymmetric: allowing workers to use different local batch sizes produces 8.7× higher training throughput in the reported imbalance case than forcing an identical size on both GPUs.

Milliseconds are the enabling result

The most transferable measurement is adjustment latency. Waiting for a batch boundary can exceed 250 ms on one GPU and one second when several GPUs synchronize. SIRIUS adjusts training memory in less than 5 ms on average, 121× faster than the naive handover, with reported P99 values of 11.4 ms on one GPU and 10.5 ms on four. Its allocation path averages 0.8 ms including zero-filling, and the complete reallocation is 148× faster than the naive path.

That time-scale change alters the control policy. If reclaiming memory takes seconds, the scheduler must predict far ahead and reserve for peaks. If it takes milliseconds, the scheduler can respond to an observed arrival and return capacity after the burst. Prediction is still useful for preloading, but a forecast error no longer means a full cold start in every case.

The paper also checks whether dynamic batch sizes damage training. Swin-T reaches the same 75% CIFAR-100 accuracy in an average 210.6 epochs with the dynamic scheme versus 206.2 without it, across five runs. This is not a general convergence proof, but it prevents the throughput result from quietly trading away the stated training target.

The LLM case changes the memory object

For LLM inference, model weights remain pinned while the KV cache grows and shrinks with requests[3]. SIRIUS shares the remaining space with Qwen2-0.5B training while serving Llama2-13B through vLLM on an A100 80 GB. The arrival trace comes from BurstGPT, scaled to ten requests per second. Time to first token and time between tokens are each given SLOs at four times standalone latency.

Compared with a 50% static split, SIRIUS raises TTFT compliance by 40 percentage points and TBT compliance by 7 points. Compared with a 75% inference reserve, it improves training throughput by 1.5×. Against inference-only, it preserves 89% and 91% of TTFT and TBT compliance, respectively. The result is smaller than some image-model tests because an LLM’s pinned weights and growing KV cache reduce the truly elastic part of memory.

This distinction should guide deployment. SIRIUS can lend KV capacity to inference, but it cannot make model weights disappear. Multi-model LLM services with frequent weight loading may still need a model-residency scheduler or disaggregated memory tier. The handover mechanism solves rapid capacity adjustment inside one device; it does not solve every placement problem above it.

Where the current design stops

The training elasticity evaluated in depth is data parallelism. Workers can change local batch size and still combine gradients. Pipeline-parallel or tensor-parallel training partitions model state and execution structure across devices. Returning memory from one stage can require moving layers or reshaping tensors, which is much slower than resizing a batch. The authors identify resharding as future work rather than including it in the reported fast path.

Interference outside memory also remains. Inference is allowed to use compute without restriction, so training throughput can collapse during sustained heavy load. PCIe, CPU preprocessing and framework control paths can become shared bottlenecks even when HBM allocation is correct. A production policy needs a minimum progress guarantee for training or must accept that it is opportunistic work.

Failure handling needs explicit design. If a training worker does not reach its adjustment point before an inference deadline, the controller must choose between rejecting the request, switching devices or terminating work. The published prototype demonstrates the normal handover path, not a complete multi-node failure protocol. Operators should test stalled kernels, GPU errors and control-plane delay before putting latency-critical traffic on the same device.

The accounting that matters

Average GPU utilization is a poor success metric for this system. A training job can drive utilization to 100% while causing every inference request to miss. The correct numerator is useful background work; the constraints are TTFT, TBT or request-latency compliance for the foreground service. Reports should show completed training samples per accelerator-hour at fixed inference arrival traces and fixed SLO attainment.

Memory telemetry must also distinguish weights, KV cache, activations, optimizer state and temporary workspace. Their lifetimes and reclaim costs differ. A single used-memory gauge cannot tell the controller which bytes can move within 5 ms. SIRIUS works because the training runtime exposes semantic allocation boundaries that the device driver alone cannot infer.

Finally, the reserve watermark should be charged as insurance. Keeping 10 GB free is not “unused” if it prevents a cold start with a known probability. The operating question is whether the expected SLO loss avoided by that reserve is worth more than the training work it displaces. That calculation will differ by customer tier and time of day.

The system decision

SIRIUS turns GPU colocation from a fixed partition into a priority relationship. Inference may claim the device when demand arrives; training absorbs the variance by changing its batch and memory footprint. The sub-5 ms average adjustment is the result that makes this relationship operationally plausible.

The approach is most attractive for bursty inference paired with data-parallel training that can tolerate changing local batches. It is less ready for tightly coupled model-parallel jobs or sustained inference saturation. An infrastructure team should adopt it only after measuring adjustment tails, convergence, foreground SLOs and background progress together. When those four remain healthy, idle HBM becomes recoverable capacity rather than a permanent reservation.

This article is an independent editorial digest of the USENIX ATC 2025 paper[1]. The prose and figure were created anew for Silicon & Systems; no paper figure or table was reproduced. Measurements are reported with the authors’ workload and SLO definitions. Copyright in the original paper remains with its authors and the USENIX proceedings (2025).