An inference cluster is sized for bursts and latency objectives. A fine-tuning cluster is sized for sustained progress. Keeping them separate simplifies admission control, but it also leaves stranded capacity whenever online traffic falls or a training queue is temporarily empty. Moving whole GPUs between pools is too slow and coarse for request-level variation. FlexLLM asks whether the two workloads can share a model and a GPU iteration without making the online user wait.

The answer depends on more than priority scheduling. Fine-tuning carries forward activations, backward gradients, and optimizer work that inference does not. A training step can monopolize memory or create a long non-preemptible kernel even if the scheduler intends to favor requests. FlexLLM combines compile-time reduction of the parameter-efficient fine-tuning graph with a runtime that represents training progress in token-sized pieces. Inference keeps the deadline; training consumes the safe remainder.

Why allocating whole GPUs is too coarse

Modern services often host many adapters over one frozen base model. Parameter-efficient fine-tuning updates a small set of weights, but conventional training stacks still allocate activation memory and launch a graph designed for batch throughput. Running that stack beside an inference server duplicates base-model storage or forces the two systems to coordinate at process boundaries. Static partitioning protects one workload by reserving resources the other cannot use.

Demand also changes on different time scales. Inference varies by seconds, while fine-tuning jobs run for minutes or hours. A cluster manager can move a job after a sustained shift, but it cannot reclaim the memory and compute gaps between successive decode iterations. Co-serving must therefore operate below the job and batch level. It needs a unit small enough to yield before the next latency-sensitive request becomes late.

FlexLLM uses tokens as that unit. The scheduler chooses inference tokens and fine-tuning tokens for the next co-serving iteration. It can reduce or suspend training work as request load rises, then restore it when latency headroom returns. This is not arbitrary kernel preemption; the training graph and its state are arranged so useful progress can be divided at those boundaries.

FlexLLM’s co-serving mechanism and evidence boundary. Inference and fine-tuning tokens occupy one iteration under SLO feedback instead of reserving separate GPU pools. The paper reports 1.9x to 6.8x higher fine-tuning throughput across load regimes while maintaining tested inference objectives up to 20 requests per second. Original figure created for this article.

Making the training graph fit beside inference

Memory is the first constraint. FlexLLM uses dependent parallelization and graph pruning to remove activation state that does not need to survive for parameter-efficient updates. The paper reports end-to-end GPU-memory savings of up to 80% for the fine-tuning path. This capacity becomes room for inference KV cache, larger request batches, or additional adapter training.

Dependent parallelization recognizes relationships among operations instead of applying one parallel strategy independently to every layer. Graph pruning removes computation and stored intermediates that cannot affect the trainable parameters. The base weights remain shared and frozen. These compile-time choices are what make token-level runtime scheduling credible; a scheduler cannot reclaim memory that the graph permanently reserves.

Fine-tuning tokens still require forward and backward work. FlexLLM tracks progress through the microbatch and interleaves it with inference rather than treating one full training step as indivisible. The runtime also has to preserve optimizer semantics and gradient accumulation. Pausing training changes wall-clock progress, not the mathematical order of updates within the supported PEFT configurations.

A hybrid scheduler with two outputs

A normal inference scheduler optimizes request metrics such as time to first token, time between tokens, or deadline attainment. A training scheduler optimizes samples or tokens processed. FlexLLM must produce both. It admits latency-sensitive inference work first, estimates the remaining budget in an iteration, and fills that budget with fine-tuning tokens. Feedback from observed latency changes the mix as load varies.

The training objective is not instantaneous utilization. A GPU can appear fully busy while long training kernels cause repeated request misses. Conversely, leaving a small amount of capacity unused may protect a tail-latency target and produce more valid requests over time. The correct metric is inference work completed within the SLO together with fine-tuning progress, each measured separately.

Token granularity also has overhead. Smaller pieces improve responsiveness but increase scheduling and launch costs. The system relies on static graph optimization and batched execution to keep the useful units large enough for the GPU. Workloads with very short kernels, unsupported training operators, or full-parameter updates may expose a different balance.

What the evaluation shows

The end-to-end study uses LLaMA-3.1-8B, Qwen-2.5-14B, and Qwen-2.5-32B. Under the tested inference service-level objectives, FlexLLM maintains compliance at offered loads up to 20 requests per second. Compared with isolated or competing co-serving allocations, it improves fine-tuning throughput by 1.9x to 4.8x under heavy inference and by 2.5x to 6.8x under light inference. At peak demand it retains more than 76% of peak fine-tuning progress in the reported configurations.

These ranges should not be collapsed into one universal speedup. Light-load gains include capacity that an isolated serving reservation leaves unused. Heavy-load gains depend on how much slack remains after protecting inference. Model size, adapter rank, sequence lengths, KV-cache pressure, SLO definition, arrival distribution, and GPU type change the available budget. A service with no latency slack should expect training to approach zero rather than assume a fixed minimum share.

The 80% memory result is also a maximum across evaluated configurations, not an 80% reduction in total GPU memory for every deployment. Base-model weights and inference KV cache remain. The relevant procurement number is how many SLO-valid requests and how much fine-tuning progress fit on a given GPU under the actual model and context distribution.

Isolation moves from hardware to policy

Sharing GPUs removes the strong failure and performance boundary created by separate pools. A fine-tuning input with an unexpected sequence length can increase memory pressure. A compiled training operator can run longer than its profile. An inference burst can starve training for an extended period. The scheduler therefore becomes part of the service reliability boundary.

Operators need explicit minimum and maximum shares, out-of-memory containment, cancellation, and accounting per tenant. Training progress should be checkpointed because it may advance irregularly. Inference admission control must consider training state already resident on the device. A policy that simply suspends fine-tuning after latency degrades reacts too late if a long kernel has already launched.

Security matters as well. Sharing a base model and device memory across adapters can reduce capacity cost, but it increases the need for memory clearing, adapter ownership checks, and isolation of gradients and training examples. FlexLLM’s performance design does not remove those platform responsibilities.

Capacity planning with a two-dimensional curve

Traditional planning produces one throughput curve for inference and another for training. Co-serving needs a frontier. Each point states an inference arrival rate and SLO attainment together with the resulting fine-tuning progress. The useful operating point is not necessarily the highest combined token count because inference and training tokens have different value and cost.

The curve should be measured across diurnal traffic, prompt and output lengths, adapter ranks, and model sizes. It should include p95 or p99 latency, not only average compliance. Teams should also record the transition time when load rises, since a scheduler that converges after several missed deadlines may be unsuitable even if steady-state metrics look good.

Power can add a third constraint. Filling inference gaps with training raises energy use compared with leaving GPUs idle. This may improve cost per completed task while violating a rack power cap or increasing cooling demand. The scheduler can only make the right decision if it receives the facility limit and the relative value of training progress.

Where the design applies

FlexLLM is best matched to services that own both inference and PEFT workloads for the same base-model family. Shared weights and a queue of deferrable fine-tuning work create the opportunity. Full pre-training, full-parameter fine-tuning, unrelated model architectures, or strict tenant isolation reduce it. Training that must finish by a deadline also cannot be treated as unlimited background work.

The design can complement cluster-level reallocation. Token scheduling handles seconds-scale variation, while the cluster manager changes the number of co-serving replicas when the demand regime persists. Without the outer loop, a heavily loaded fleet may keep fine-tuning resident but make negligible progress. Without the inner loop, whole-GPU movement leaves short-lived gaps unused.

An adoption test should therefore include both control loops, checkpoint recovery, unexpected bursts, and a training completion objective. Measuring only a stable trace risks selecting a policy that performs well before the first real traffic shift.

The system decision

FlexLLM reframes spare GPU capacity as time and memory inside an iteration, not as an idle whole device. Compile-time graph reduction makes PEFT small enough to coexist, and token-level scheduling spends only the budget left after inference objectives are protected.

For infrastructure teams, the important contract is asymmetric. Inference owns a measurable deadline; fine-tuning receives elastic progress and must tolerate pauses. If both workloads require hard completion guarantees, static separation may remain simpler. If training is deferrable and shares a base model, the reported gains show that isolation by whole GPU can be unnecessarily expensive. The buying metric then becomes a frontier of SLO-valid inference and useful training progress, not a single utilization percentage.

This article is an independent editorial digest of the NSDI 2026 paper[1]. The prose and figure were created anew for Silicon & Systems; no paper figure or table was reproduced. Results retain the authors’ models, PEFT configurations, traffic, SLOs, and baselines. Copyright in the original work remains with its authors and the USENIX proceedings (2026).