GPU clouds already pay for two communication planes. The scale-up fabric connects accelerators inside a server or tightly coupled domain. Ethernet or InfiniBand connects hosts and racks. Collective libraries normally select the first plane for traffic among GPUs that share it, leaving the second path outside that transfer even when its NIC and switches have spare capacity. MPCCS asks whether one collective can use both at once.
The answer is not to hash packets over two links. Scale-up and scale-out paths differ in bandwidth, latency, ordering, topology, and the amount of state visible to software. Modern interfaces also offload transport into hardware, so a runtime cannot observe every acknowledgement or freely move a live flow. Collective operations add another constraint: the slow completion of one dependent flow can delay every rank. MPCCS therefore coordinates windows and coflows above the hardware transports.
Two fabrics, one completion time
All-reduce, all-gather, reduce-scatter, and all-to-all create sets of dependent transfers. A ring all-reduce, for example, advances through repeated reduce-scatter and all-gather steps. Sending some bytes over a second network helps only if the partitions complete at compatible times. A poor split moves the bottleneck rather than removing it, and excessive reordering can consume buffers or stall the collective schedule.
The scale-up path is usually faster, but it is not infinitely fast and may be shared by other GPUs or collectives. The scale-out path has a longer round trip yet can carry an independent portion. Its contribution depends on the bandwidth-delay product, congestion, message size, and the collective’s dependency graph. A fixed 80:20 ratio cannot remain optimal as either network changes.
MPCCS exposes two logical transmission windows to the collective runtime. One window feeds the scale-up interface and the other feeds the scale-out interface. Traffic is divided at a boundary the offloaded hardware can accept, without requiring the NIC or accelerator fabric to expose its internal transport state. The runtime advances the windows while preserving the operation’s byte ranges and completion semantics.

Estimating capacity hidden by offload
Software multipath transports often use observed round-trip time and delivered throughput to update a split. Hardware offload hides much of that state. MPCCS estimates the effective bandwidth-delay product from the progress visible at its windows. The estimate determines how much data each path should keep in flight, then adapts when contention or message characteristics change.
Bandwidth alone would be insufficient. A high-bandwidth path with a longer feedback delay needs a larger window to stay occupied. A low-latency path may finish a small message before the second path contributes. The controller must therefore consider both rate and delay, and it must avoid oscillating as two independent networks respond to load.
This design also explains why the second network is not always useful. Short collectives can end before setup and synchronization costs amortize. A congested scale-out network may contribute negative value. A topology in which the NIC path crosses additional PCIe copies can lose the supposed parallelism. MPCCS needs a policy that can assign nearly all traffic back to one path when measurement says so.
Coflow synchronization preserves collective progress
A collective contains several flows whose usefulness is tied to a common stage. Optimizing each independently can leave one laggard on the second network and hold the next stage. MPCCS groups related transmissions as a coflow and synchronizes their multipath decisions. The objective is the completion time of the collective stage, not the sum of per-flow throughput.
The paper argues that this mechanism applies across diverse collective shapes. Correctness requires each byte range to arrive once and the collective’s reduction and ordering rules to remain intact. Performance requires partitions to finish together closely enough that the next dependency is not delayed. The coflow view provides the common control unit while individual windows carry the bytes.
Collective libraries and frameworks still decide the algorithm and rank topology. MPCCS adds another dimension to that plan. An algorithm optimized for one scale-up hierarchy may no longer be best when scale-out capacity is available inside the same stage. A production integration should therefore search algorithm, channel count, and path split together rather than attach multipath after the schedule is fixed.
Reading the measured gains correctly
The reported communication microbenchmarks improve by 23% to 54% over vanilla NCCL. End-to-end LLM training accelerates by more than 10% in the evaluated configurations. The smaller application gain is expected because compute, input, optimizer, and communication that cannot overlap or use both paths remain. It also indicates the right denominator: reduced collective time must be translated into completed training steps.
The study combines a testbed with simulation to cover hardware and scale conditions. That broadens the topology space but requires separating measured and simulated results. The maximum microbenchmark speedup should not be used as a rack-wide training forecast. Model parallelism, tensor size, collective mix, compute-to-communication ratio, network contention, and the placement of jobs determine the end-to-end outcome.
The baseline matters as well. NCCL versions and topology discovery improve, and vendors can select multiple rails within a network class. MPCCS’s distinct claim is simultaneous use of scale-up and scale-out planes, not that every default NCCL configuration is permanently 54% slower. Reproduction should pin library, driver, firmware, algorithm, and channel settings.
Failure, congestion, and ownership
Using two paths expands the failure surface. If one plane resets, the runtime must know whether a window’s bytes were delivered and whether the collective can continue on the remaining path. Hardware-offloaded transports may not expose enough state for transparent migration. A safe implementation can fail the operation and retry from a known boundary rather than infer completion.
Congestion isolation also changes. Scale-out Ethernet may serve storage, tenant traffic, or other distributed jobs. Borrowing its capacity for an intra-domain collective can harm workloads that were planned independently. Quality-of-service classes and admission control must price the extra traffic. The best local training time is not necessarily the best cluster-wide goodput.
Ownership spans multiple teams. Accelerator runtime engineers control collective semantics; network teams control routing and congestion; cloud schedulers place jobs. MPCCS needs telemetry and rollback across all three. Without a shared incident view, a path split that adapts to network load may look like an unexplained application change.
A procurement test for dual-plane systems
Hardware specifications should expose the simultaneous-use constraint. Some architectures advertise high scale-up and scale-out bandwidth but share PCIe switches, DMA engines, memory controllers, or power limits. Summing link rates can overstate usable concurrency. A test should run both paths together with the actual collective and observe delivered bytes, GPU copy engines, host memory traffic, and application progress.
The test matrix should include small and large messages, each collective type, changing background traffic, and partial link failure. It should record the controller’s convergence time and the distribution of path shares, not only the best steady-state throughput. A path that helps large all-reduce but hurts latency-sensitive all-to-all may need workload-specific admission.
Cost analysis should count software and operational complexity. A second path that adds 10% training throughput can be valuable on a large fleet, but only if debugging, congestion policy, and failure recovery remain manageable. Conversely, an already provisioned scale-out network may provide cheaper incremental capacity than upgrading every scale-up domain.
The system decision
MPCCS demonstrates that network hierarchy need not dictate exclusive use. A scale-out path can complement a scale-up collective when the runtime controls independent windows, estimates effective capacity, and synchronizes the flows that determine stage completion.
The result changes how a GPU cloud should describe its fabric. Two link specifications are not two additive resources until software can schedule them jointly under real contention and failure. The relevant metric is SLO- or step-valid collective capacity delivered by both planes together. If that control exists, an idle NIC becomes useful accelerator bandwidth. If it does not, the second network remains a separate asset whose headline rate cannot shorten the collective.
Source and copyright note
This article is an independent editorial digest of the EuroSys 2026 paper[1]. The prose and figure were created anew for Silicon & Systems; no paper figure or table was reproduced. Testbed measurements and simulated results are distinguished in the analysis. Copyright in the original paper is held by its authors and publication rights are licensed to ACM (2026); the paper is distributed under CC BY-NC-ND 4.0.