Cloud operators want virtual machines because they provide an established boundary for provisioning, images, access control, accounting, and recovery. Large-model training wants a short path from GPU memory to the network. A conventional virtual NIC can insert host copies and software switching between those goals. Vela shows that the boundary can remain a VM while the data path bypasses most of the host.

The ASPLOS 2025 paper documents an A100-based system operated by IBM for about 2.5 years. KVM supplies the guest boundary. GPUs and SR-IOV virtual functions expose devices directly to the guest. GPU Direct RDMA lets the NIC move data to and from GPU memory, and RoCE carries that traffic across Ethernet. The architecture uses commercial servers and networking rather than a proprietary scale-up machine, so its contribution is the integration and operating experience.

The data path that virtualization must preserve

Without direct assignment, a GPU collective can cross device memory, host memory, a virtual switch, a host network stack, and a physical NIC. Copies, interrupts, CPU scheduling, and address translations add latency and consume bandwidth. They also make performance depend on unrelated host activity. Distributed training repeats this path for collectives at every step, so a modest per-transfer cost becomes synchronized GPU idle time.

SR-IOV divides a physical NIC into virtual functions that can be assigned to guests. IOMMU and device assignment isolate DMA addresses while avoiding a software switch for the main data path. GPU Direct RDMA extends the path to GPU memory. The guest still runs a virtualized control environment, but payload bytes travel between the GPU and NIC without staging through a host buffer.

RoCE adds a network requirement. RDMA assumes low loss and predictable congestion behavior, while ordinary Ethernet can drop packets under incast. Priority flow control, ECN, routing, buffer policy, and queue configuration must be engineered as a fabric. A direct NIC is not sufficient if the network turns collective bursts into congestion pauses.

Vela’s mechanism and evidence boundary. A KVM guest owns assigned GPUs and an SR-IOV NIC function; GPU Direct RDMA keeps payloads off the host-copy path and RoCE carries collectives across Ethernet. The paper reports results at roughly 1,500 A100 GPUs, including about 80% of ideal model-training throughput. Original figure created for this article.

Hardware placement becomes a cloud API

Device assignment exposes physical topology to the scheduler. GPUs, NICs, PCIe switches, CPU sockets, and NUMA memory are not interchangeable. A VM that receives eight GPUs but a NIC behind the wrong root complex can cross CPU interconnects before reaching the network. The cloud control plane must allocate a topology-valid bundle rather than count devices independently.

Vela’s design therefore joins provisioning with performance. Images and lifecycle remain cloud-native, while placement respects GPU-to-NIC affinity and network rails. Firmware, driver, CUDA, OFED, and guest-kernel combinations must also be qualified together. A VM snapshot or migration mechanism designed for ordinary servers cannot assume it can move a passed-through GPU and RDMA queue transparently.

The physical service unit remains a server tray and rack even when users see VMs. Cable failures, NIC resets, switch maintenance, and thermal limits occur below the guest. Operators need inventory that maps a training rank back to those components without exposing another tenant’s resources.

Conceptual rack-scale material view for Vela. The generated material layer shows GPU server trays, NIC regions, disciplined cabling, and a top-of-rack Ethernet switch; deterministic labels identify the GPU Direct and SR-IOV path. It is not an IBM product photograph, rack floorplan, or manufacturing drawing. Original figure created for this article.

What the scaling results establish

At roughly 1,500 GPUs, Vela reaches about 80% of ideal throughput while training a 50-billion-parameter decoder model with model parallelism. High-Performance Linpack reaches about 70% of the per-GPU FLOPS observed in a single VM. The two percentages answer different questions. Model training includes the real communication pattern and software stack; HPL stresses a regular numerical workload and provides a narrower scaling reference.

The paper and related IBM account also compare TCP, RoCE, and GPU Direct RoCE paths. IBM reports that enabling GPU Direct RDMA over Ethernet improved network throughput by two to four times and reduced latency by six to ten times in its upgrade context. Those are before-and-after results for a particular implementation, not general ratios between all TCP and RoCE clusters.

Approximately 80% scaling means the system loses about one fifth of the ideal linear gain at that model, GPU count, and parallel plan. It does not imply 80% average GPU utilization across the fleet or 80% application availability. The result is nevertheless important because the guest boundary remains in the measured path. Virtualization did not force the job onto a conventional virtual network stack.

Operational lessons from 2.5 years

The long deployment period covers problems that a short benchmark misses. Driver and firmware changes can alter peer-to-peer access. A host update can reset device assignment. RoCE behavior can change with a switch configuration intended for another traffic class. A failed component can leave a job alive but slow, and the VM boundary can make host and guest telemetry appear in separate systems.

Cloud operations need coordinated health states. The scheduler should avoid a host when GPU, NIC, PCIe, or fabric telemetry shows degradation, not only when the VM agent reports failure. Training frameworks need rank-level failure information, while infrastructure teams need the physical mapping. Recovery should preserve an immutable VM image and checkpoint while permitting replacement hardware to differ.

Maintenance is another tradeoff. Bare-metal clusters can expose custom diagnostics and firmware directly; virtualized clusters can roll tested images and isolate users from host changes. Device passthrough reduces live-migration flexibility, so Vela’s kind of cloud system gains reproducibility and provisioning speed rather than ordinary VM mobility.

Isolation is broader than the hypervisor

KVM isolates CPU and memory mappings, and the IOMMU constrains DMA. Shared switches, NIC queues, PCIe links, power, and cooling remain contention points. A neighboring job can affect collective latency without violating memory isolation. AI cloud service-level objectives therefore need performance isolation and telemetry across the physical path.

RoCE configuration can also create coupled failure. Priority flow control can propagate pauses if congestion is not bounded. ECN thresholds and congestion control must match switch buffers and traffic patterns. Multi-rail routing needs stable entropy without reordering RDMA flows. These controls sit outside the guest but determine its effective communication performance.

Security review must include NIC firmware, virtual-function reset, IOMMU groups, and peer-memory registration. Granting a guest direct access improves performance by removing host mediation; it increases the importance of hardware isolation and tested driver paths. A performance result cannot substitute for that review.

A fair comparison with bare metal

The correct experiment holds hardware, model, parallelism, software versions, and network configuration constant, then compares the VM path with a qualified bare-metal path. Kernel microbenchmarks alone are insufficient. The test should include collective bandwidth by message size, tail latency under competing traffic, complete training throughput, failure recovery, image deployment, and maintenance time.

Cost should include utilization outside one job. Virtualization can reduce turnaround by offering consistent images and multi-tenant allocation. Device passthrough can create fragmentation when GPU and NIC bundles no longer fit pending requests. A bare-metal cluster may deliver a slightly faster job but lower fleet utilization, or the reverse. The useful denominator is completed training work per rack and per operator hour under availability targets.

New accelerators require repeating the qualification. A100 results do not establish H100, H200, or future GPU behavior, whose NVLink, PCIe, NIC, and confidential-computing features differ. The lasting method is to preserve the direct path and verify topology rather than assume one configuration transfers unchanged.

The system decision

Vela demonstrates that virtualization overhead is not a single percentage attached to a VM. It is the result of a specific path. When a guest’s collective traverses host buffers and software switching, the penalty can dominate. When GPUs and an SR-IOV NIC are assigned coherently and GPU Direct reaches a tuned RoCE fabric, the cloud boundary can coexist with large-scale training.

The decision for an operator is whether it can own the full bundle: topology-aware placement, device and firmware qualification, congestion control, rank-to-hardware telemetry, checkpoint recovery, and security. If those systems are missing, passthrough merely hides unmanaged bare metal behind a VM. If they are present, virtualization becomes an operational interface around a near-direct GPU network path, which is the real lesson behind Vela’s 1,500-GPU result.

This article is an independent editorial digest of the ASPLOS 2025 paper[1] and related IBM technical materials[2][3]. The prose and figures were created anew for Silicon & Systems; no paper figure or table was reproduced. The rack material view is conceptual and not a product image. Copyright in the original paper is held by its authors and publication rights are licensed to ACM (2025).