Alibaba’s ASI platform exposes 155,410 GPUs to 81 internal departments, which sounds like a pool large enough to make fragmentation disappear. A six-month production trace shows the opposite[1]. More than 99% of jobs request an exact GPU model, large jobs also need contiguous network topology, and a free accelerator can remain unusable because its host lacks the requested CPUs. At one point, jobs asking for eight GPUs and 128 CPU cores could not start while as much as 28% of the relevant GPUs sat idle.
The OSDI 2026 paper is valuable because it describes the shared fleet as an operating system rather than an inventory. Heterogeneous hardware, high-priority services, overnight slack and application portability are all scheduling state. Alibaba combines placement repair, topology-aware packing, reclaimable workloads and device-specific software optimization. GPU allocation is 68% when only high-priority work is counted, then reaches 93% after low-priority work fills the gaps. The difference between those numbers is not more silicon. It is a more accurate definition of available.
One fleet, several incompatible shapes
The workload mix differs from earlier public traces[2][3]. Online inference accounts for 54.6% of jobs, offline inference for 21.6%, training for 19.5% and development for 3.4%. Generative AI dominates training at 71% and offline inference at 98%, while recommendation models still represent 63% of online inference. The median execution time is five hours, compared with 23 minutes in an earlier Alibaba PAI trace and two minutes in Microsoft’s Acme trace. The scheduler therefore manages long-lived placements whose mistakes persist.
Priority divides the fleet into two markets. Online inference and most training are high priority, while offline inference fills the low-priority side. Development is also treated as high priority because an engineer waiting for a GPU is part of the service objective. Median scheduling latency is only one second, yet the high-priority 90th percentile reaches 101 seconds. Task duration is much longer, with a 90th percentile near 24 hours, so even modest fragmentation accumulates around placements.
Hardware diversity does not automatically create application flexibility. Fewer than 1% of jobs span different GPU models, and more than 99% pin one model exactly. The fat-tree network adds another shape constraint. One access switch connects 32 to 64 eight-GPU nodes, or 256 to 512 GPUs, and all-reduce within one access-switch domain performs up to 27% better than communication that crosses domains. A request for 256 GPUs is therefore not satisfied by any 256 free devices. It wants one model, sufficient host CPUs and preferably one network neighborhood.

Repair the placement without stopping the service
Alibaba’s in-place compaction (IPC) repairs fragmentation by moving existing allocations. A global integer program would find better arrangements but does not finish at this scale. The paper reports that an optimal solver can take five minutes for 25 nodes and two days for 100, while a production partition may contain 500 nodes. IPC instead divides the cluster, applies a recursive ejection chain with depth three and repeats at most five times.
An ejection chain moves the allocation occupying a desired node, then finds a new home for whatever that move displaces. The mechanism is make-before-break: it starts the replacement container before releasing the original, thereby preserving high-priority service. Decisions complete in under two minutes and reduce partially occupied nodes by 20.2%. A separate entropy-based placement policy packs large distributed jobs into fewer access-switch domains, preserving the remaining contiguous regions for future requests.
This repair exposes an important change in fragmentation. Earlier GPU clusters devoted substantial attention to fractional sharing, but the ASI trace finds it negligible. Generative workloads consume whole devices, and the dominant losses now come from stranded GPUs on partly occupied nodes, insufficient CPU cores and network topology. Scheduling policy has to follow the workload rather than continue optimizing yesterday’s bottleneck.
Sell the slack twice, but preserve the first buyer
High-priority demand leaves a strong daily cycle, including as many as 10,000 standby GPU-hours around midnight. Alibaba recovers that capacity through two mechanisms. The first keeps online-inference containers warm but detaches traffic and places them into a sleep state. When demand returns, eviction takes 13 seconds on average and 48 seconds at the 95th percentile; fewer than 5% require a forced termination. Low-priority work harvests 90% of the standby GPU-hours without paying the cold-start cost of rebuilding the service.
SpotGPU handles general preemption. Instead of evicting the smallest or newest job, it estimates lost work as the number of GPUs multiplied by the time since the last checkpoint. The scheduler then takes capacity from jobs with the lowest recomputation cost. In production experiments, this policy reduces low-priority completion time by 24% without measurable damage to high-priority performance. The result is the 68% to 93% allocation step reported for combining both priority classes.
Portability remains the harder form of slack recovery. One alternative accelerator, identified as XPU-A, has a theoretical advantage over H20 for a target workload but initially delivers only 80% of H20 performance. Kernel and framework work improves its throughput by as much as 43%, including a 1.58× prefill-kernel gain, after which high-priority demand for the device grows by 2.5×. A heterogeneous fleet therefore needs a software budget for every hardware option. Otherwise, the scheduler sees capacity that applications will not request.

Utilization is not one number
The remaining telemetry warns against treating allocation as useful computation. Online inference has median streaming-multiprocessor utilization of only 6% and memory use of 30%. Generative inference reaches 94% memory use but only 5% SM use, while conventional deep neural networks show about 20% and 6%, respectively. These complementary profiles suggest colocation, but performance isolation and memory pressure make it harder than adding two percentages.
CPU colocation creates a less visible interference path. CPU-only jobs dominate CPU use on 80% of GPU nodes. Their presence reduces median training SM utilization by 10% and the 90th percentile by 18%. A scheduler can therefore report a highly allocated GPU fleet while co-located CPU work quietly lowers completed training tokens.
The paper explicitly excludes giant foundation-model pretraining on dedicated clusters, so its findings should not be applied unchanged to every hyperscale system. It instead describes the broad middle of a multi-tenant AI cloud, where inference, development and training compete across several accelerator generations. For that environment, the durable lesson is an accounting rule. Physical idle, schedulable idle and productive idle are different quantities. A neocloud that sells only the first number will discover the other two in its queue.
A utilization report needs three denominators
The 68% and 93% figures describe allocated devices, not useful model work. The paper’s own telemetry shows why the distinction matters: an inference GPU can be almost full in memory while using only a small fraction of its arithmetic units, and CPU neighbors can lower training SM utilization even when the GPU remains assigned. A fleet report should therefore separate physical utilization, schedulable utilization, and productive utilization. The first asks whether a device is allocated, the second whether idle capacity can satisfy an actual request shape, and the third whether the allocated device advances accepted work at the expected rate.
These denominators expose different remedies. Physical idle invites more demand. Schedulable idle requires compaction, topology repair, CPU balancing, or software portability. Low productive use calls for batching, colocation, kernel work, or a different accelerator. Collapsing them into one percentage encourages the wrong investment, such as buying GPUs to solve a CPU-fragmentation problem or forcing colocation onto a memory-bound service with no isolation margin.
For a neocloud, the customer-facing implication is equally direct. Availability should be quoted for a requested shape, not for the fleet in aggregate: GPU model, count, CPU and memory ratio, network locality, start deadline, and expected duration. The operator should then report the probability that this shape starts on time and the goodput it sustains after placement. Alibaba’s trace demonstrates that scale does not erase fragmentation. It multiplies the number of resource dimensions that must align. The competitive advantage belongs to the scheduler that can translate nominal inventory into predictable job starts without hiding the cost in lower runtime efficiency.
Idle capacity must pass a schedulability test
A GPU is not available merely because its utilization counter is low. It may belong to a partially occupied topology, hold state for a latency-sensitive service, lack the memory or interconnect required by a waiting job, or be expected to return to its primary tenant before a new task can finish. ASI’s contribution is easier to understand when availability is treated as a predicate rather than a percentage.
The predicate should include device type, memory, host and fabric locality, health state, software image, expected free interval, preemption cost, and the job’s minimum gang size. A scheduler can then distinguish immediately usable capacity from reclaimable capacity and stranded fragments. Reporting all three prevents a fleet from claiming idle GPUs that no queued job can assemble into a valid placement.
Time is as important as shape. A two-hour gap may fit evaluation or inference work but not a training stage whose setup and checkpoint cost consume most of the interval. The platform needs duration estimates for both the borrowed capacity and the job. Admission should include a safety margin for forecast error and a recovery path when the primary service returns early.
Migration must pay back before the placement changes again
Repairing a fragmented placement can improve locality and release complete groups of devices, but moving a running service consumes network, host memory, and warmup time. The benefit must be measured over the expected lifetime of the new placement. A short-lived compaction that saves a few GPUs can lose more work than it creates.
The migration ledger should include bytes of weights and runtime state, traffic on each network tier, pause or overlap time, cache warmup, performance after the move, and any effect on neighboring jobs. It should also record the capacity that became schedulable, not only the number of moved instances. Freeing seven isolated GPUs may help no waiting gang, while freeing one connected eight-GPU island can admit a job.
Rollbacks need equal design. If the destination is slower or a health signal changes, the service must return without duplicating ownership or losing in-flight requests. Repeated movement can be limited with hysteresis and a minimum residence time. These controls turn compaction from a perpetual optimizer into a bounded operational action.
Selling slack twice requires a priority contract
Borrowing capacity for lower-priority jobs increases utilization only if the first owner can reclaim it predictably. The contract must define notice, checkpoint or eviction behavior, maximum interruption, and compensation for lost work. A low-priority tenant should know whether it is buying a full GPU-hour, a preemptible interval, or an opportunistic token budget.
The primary tenant also needs protection from hidden interference. Host CPU, memory bandwidth, storage, and fabric can remain shared even when GPUs are partitioned. A borrowed job that saturates one of these resources may violate the service that owns the devices. Isolation tests should cover the complete node and network path, not only accelerator memory.
Fairness cannot be reduced to priority order. A stream of short high-priority arrivals can prevent a large job from ever assembling its gang. Reservations, aging, or bounded admission windows may be needed so the scheduler does not maximize immediate utilization while starving work that produces more completed value over time.
Fleet reports need three denominators and one outcome
Installed GPUs describe capital. Healthy and powered GPUs describe operational supply. Schedulable GPUs describe supply that can satisfy the active job shapes and time windows. Busy GPUs describe allocation, not necessarily useful work. These numbers answer different questions and should be shown together.
The outcome is completed workload: training steps, evaluated checkpoints, or inference requests within their SLO, normalized by installed accelerator-hours and cost. This measure captures fragmentation, migration, preemption, failures, and interference without pretending they are the same mechanism. It also reveals when higher busy time comes from low-value background work that slows the primary fleet.
We read ASI’s 155,410-GPU evidence as a call for auditable capacity accounting. Heterogeneity and service dynamics make a single utilization number less meaningful as fleets grow. The scheduler creates value by converting healthy but unusable fragments and uncertain free intervals into placements with explicit priority and recovery rules. The most important number is not how many GPUs appeared busy. It is how much additional work finished without breaking the promises that made the capacity borrowable.
Source and attribution
This article is an editorial summary prepared for Silicon & Systems. It restates the cited paper in our own words. No text, tables or figures from the paper are reproduced; both figures were created for this article from reported results. The paper is openly available in the OSDI 2026 proceedings. Copyright (c) 2026 the authors. The source is available from the USENIX paper page.