Multi-GPU kernels are usually assembled in two layers. A compiler partitions the tensor computation, then a collective library performs all-gather, reduce-scatter or all-to-all operations between otherwise local kernels. This division is convenient, but it makes communication a barrier-shaped event. Every device tends to hold the same shared input, enter the same collective and resume computation at the same logical time. The SOSP 2025 Mercury paper argues that long-context LLM operators have outgrown that boundary[1].
Mercury treats another GPU’s HBM as an explicitly managed extension of the local memory hierarchy. The compiler can keep a shared buffer distributed, shift the order in which ranks consume its pieces, and choose whether a transfer uses a copy engine, Tensor Memory Accelerator path or ordinary loads and stores. Communication is no longer a preselected kernel placed between compute kernels. It becomes part of the loop schedule.
The hidden cost of synchronized operators
Consider a matrix operation in which every GPU owns a slice of matrix A but all ranks need matrix B. A conventional tensor-parallel schedule replicates B, runs matching loop nests and inserts a collective for partial results. Replication consumes local HBM, while the identical timeline creates bursts on the same links. More importantly, the schedule cannot trade one resource for another. It cannot decide that rank 1 should read B’s first tile from rank 0 while rank 0 has already advanced, because the collective abstraction hides individual buffer lifetimes and transfers.
Long-context attention makes the limitation visible. KV tensors can exceed one device’s HBM, grouped-query attention changes the number of independent heads available for partitioning, and intra-node NVLink is much faster than inter-node RoCE. A template that works on eight GPUs inside a server may send the wrong volume across the slower tier when expanded to two or four servers. The operator is still mathematically correct, yet its data placement produces a bad network schedule.
Mercury’s central move is temporal decoupling. It can shift the inner-loop timeline by rank, so the GPUs consume shared data in a staggered order. The same buffer remains sharded across the aggregate HBM pool and moves when a consumer needs it. That adds transfers, but it also removes full replication, permits larger local tiles and spreads traffic over time. The compiler searches whether that exchange is favorable under the actual memory and topology constraints.

CommIR makes communication transformable
Mercury lowers an operator into CommIR, a loop-based intermediate representation that records reads, writes and iteration domains. Parallelization and data placement become explicit transformations. A buffer can be replicated, sharded or shifted; a loop can be reordered or tiled; a reduction can be expressed through different communication paths. Because the representation retains dependencies, the system can test whether an asynchronous schedule is legal before generating code.
The search space would be unmanageable if every GPU instruction were included. Mercury keeps optimized local kernels as patches. It searches the distributed schedule around a known-good GEMM or attention implementation, then lowers only the communication and coordination needed to connect those pieces. This separation preserves vendor-tuned computation while allowing the compiler to change the distributed algorithm.
Candidate schedules are filtered in stages. Static analysis rejects incompatible buffer sizes and dependencies. Cost estimation narrows the space. Surviving candidates are compiled and profiled on the target machine because link contention, transport engines and kernel overlap are difficult to model precisely. The output is therefore not a portable binary schedule. It is a machine-specific plan derived from a portable representation.
The design extends across operator boundaries. Different layers may prefer different sharding layouts, and converting between them can erase an operator-level gain. Mercury includes resharding edges in the graph-level cost, allowing model parallelism and operator schedules to be selected together. That is necessary for an LLM where attention, projection and feed-forward layers put pressure on different dimensions.
What Mercury finds that templates miss
The evaluation uses H100, A100 and L4 GPUs. H100 and A100 nodes connect GPUs with NVLink; L4 uses PCIe. Nodes communicate through RoCE at 50 or 100 Gbps, while intra-node bandwidth ranges from 64 GB/s on PCIe to 900 GB/s on H100 NVLink. GPU counts vary from one to eight per node and up to four nodes. This is the important test: the topology changes enough that a single collective template should not win everywhere.
For attention, Mercury is compared with RingAttention, DeepSpeed Ulysses and USP. For GEMM patterns, it is compared with cuBLAS collectives, AsyncTP and TorchInductor-generated alternatives. Workloads use Llama-3-like multi-head and grouped-query attention, batch sizes of 1 and 16, and context lengths from 4K to 2 million tokens. The implementation uses CUDA 12.6, NCCL 2.26.2 and a PyTorch 2.8-era TorchInductor.
On H100 multi-head attention at batch 16, Mercury reaches up to 4× the performance of the compared schedules. On A100 all-gather GEMM at the same batch size, the gain reaches 1.9×. Across topology experiments that vary the number of nodes and GPUs per node, Mercury reports the lowest latency with an average 2.91× speedup. The largest advantage appears on hierarchical shapes such as 2×4 or 4×2, where the schedule must distinguish fast local links from slower remote links.
That qualification matters. On a single-level topology, such as one four-GPU node or four one-GPU nodes, handcrafted methods already align with the only link tier and the gap narrows. Mercury earns its complexity when the hardware presents choices. A compiler cannot optimize a hierarchy that does not exist.
Memory capacity is part of performance
At long sequence lengths, the objective changes from selecting the fastest schedule to finding any feasible schedule. In the eight-H100 attention test, Mercury scales from 32K to 2 million tokens. The compared baselines run out of memory at the largest point, while Mercury continues by sharding KV and output buffers more aggressively. It accepts additional communication to remain within HBM capacity.
This is why the paper’s remote-memory abstraction is more than a bandwidth optimization. It exposes a Pareto frontier between latency and memory. One schedule duplicates data and runs quickly when capacity is abundant. Another distributes buffers and pays for transfers, but enables a request that otherwise cannot run. Infrastructure software should preserve both rather than reporting only the single fastest configuration.
Mercury also shows that speedup can shrink beyond one million tokens because compute begins to dominate. Once each GPU receives enough arithmetic, eliminating communication stalls accounts for a smaller fraction of total latency. That behavior is reassuring: the system is not claiming a constant multiplier. It is identifying the region where communication scheduling governs the critical path.
The tuning bill
Search-based compilers move work from human experts to profiling infrastructure. Every device generation, driver update, topology change and kernel revision can invalidate measured costs. A schedule found on 100 Gbps RoCE may not remain optimal after congestion or a link failure; a cloud instance can expose nominally identical GPUs with different placement. Production use therefore needs a cache key that includes hardware topology, software versions and relevant model shapes, plus a fallback when the cached schedule no longer meets expectations.
Compilation time also belongs in the service economics. A stable model served for months can amortize an expensive search. A multi-tenant platform loading many adapters or rapidly changing architectures may not. The paper demonstrates multi-pass search and profiling, but it does not turn tuning latency into a fleet-wide cost model. Operators should measure time to first usable schedule, profiling GPU-hours and the number of shapes that require separate tuning.
Correctness deserves equal weight. Asynchronous remote reads introduce buffer-lifetime and synchronization hazards that a collective library normally contains. CommIR’s dependency analysis provides a foundation, and the open artifact includes numerical checks against single-device FlashAttention. Production qualification should add randomized shapes, fault injection, deterministic replay and cross-version validation. A fast distributed kernel that fails rarely is a larger operational risk than a slower collective.
What changes for system architects
Mercury weakens the assumption that collective libraries are the final abstraction for accelerator communication. NCCL remains valuable as a backend, but the semantic unit above it can be smaller than a collective and larger than an individual load. That middle layer lets the compiler schedule buffer movement according to the operator’s loop structure.
The implication for interconnect design is also concrete. Faster links are useful, but exposed hierarchy and independent transport resources can be equally important. If software can distinguish copy engines, load/store paths and local versus remote links, it can overlap traffic and place data deliberately. A flat compatibility interface can hide precisely the features that justify the hardware.
For model developers, the relevant boundary moves upward. Attention head layout, KV representation and tensor dimensions determine the compiler’s legal parallel schedules. A model architecture that reduces heads or creates irregular dimensions may save arithmetic while constraining distributed execution. Compiler feedback should therefore enter model design before the architecture is frozen.
The system decision
Mercury’s strongest result is not one 4× bar. It is the consistent advantage across hardware whose compute, memory and link ratios differ substantially. The compiler wins by refusing to decide data placement, communication and local computation in separate layers. Its advantage narrows when the topology is simple and expands when a hierarchy or memory limit creates alternatives.
An infrastructure team considering this approach should ask three questions. Does the workload repeat enough operator shapes to amortize search? Can the platform expose topology and transport engines accurately? Can schedule correctness be tested as aggressively as performance? If all three answers are yes, remote HBM becomes a useful compiler-managed tier. If not, fixed collectives remain less ambitious but easier to operate. Mercury makes that trade visible, which is more valuable than pretending that automatic optimization has no operating cost.
Source and copyright note
This article is an independent editorial digest of the SOSP 2025 paper[1] and its public artifact. The prose and figure were created anew for Silicon & Systems; no paper figure or table was reproduced. Mercury’s published article is distributed under CC BY-NC-SA 4.0, while copyright in the original work remains with its authors (2025).