An 80 GB GPU does not provide 80 GB to every operation that an LLM needs to execute. Long-context attention can create hundreds of gigabytes of intermediate state, and linear layers split across accelerators must exchange partial results before the next operator begins. Engineers usually solve those constraints by choosing a known parallel pattern, placing buffers to suit it, and manually arranging collectives around computation. The result can be excellent for one model and topology, then wrong after a change in head count, sequence length, batch size, GPU memory, or the ratio between local and network bandwidth.
Mercury, a SOSP 2025 paper from UC San Diego, Meta, George Mason University, and Rice University, changes the abstraction rather than adding another pattern[1]. Its compiler treats memory attached to another GPU as an explicitly scheduled part of the memory hierarchy. A loop-based intermediate representation called CommIR describes where computation runs, which device owns each buffer, what is replicated, and when data shifts between ranks. An autotuner then explores combinations that include familiar collective-based designs and asynchronous schedules that are difficult to construct by hand.
The reported result is broad but must be read at the right level. Across attention and matrix-multiplication operators on H100, A100, and L4 systems, Mercury averages a 1.56× speedup against the selected hand-optimized designs. In a topology sweep, it reports 2.91× average improvement, and in one H100 attention case reaches 4×. At two million tokens, competing configurations run out of memory while Mercury produces a feasible schedule by sharding more state. These are operator and one-layer measurements on testbeds of up to four nodes. They demonstrate the value of the representation and search space, not a production fleet result.
Multi-GPU operators are memory schedules
The usual description of tensor parallelism starts with arithmetic: split a matrix dimension, compute partial products, then combine them with all-gather, all-reduce, reduce-scatter, or all-to-all. That view hides two choices that often determine performance. First, a tensor can be replicated, sharded, or streamed through remote devices at different points in the loop nest. Second, communication can be synchronized as a collective or advanced in smaller units while other tiles compute.
Manual systems encode a small family of those choices. Ulysses partitions attention heads and uses all-to-all communication[4]. Ring-style attention rotates key and value blocks around devices[3]. Hybrid schemes choose between or combine these patterns according to head count and context length. Each is useful, but a template fixes important degrees of freedom before the hardware and workload are known. Grouped-query attention, for example, reduces the number of key-value heads and can remove the balance that made head-wise partitioning effective.
Mercury models an operator as loops over logical dimensions and buffers accessed by those loops. Four transformations carry the distributed meaning. parallelize maps loop iterations to devices at a specified topology level. shard assigns disjoint parts of a buffer to ranks. replicate gives each participating rank a complete copy. shift offsets access in time and rank, creating an asynchronous rotation of remote data. Loop tiling, reordering, and joining remain available from conventional tensor compilers[6]. The difference is that memory location and communication are first-class annotations in the same program.
This makes a remote buffer neither transparent shared memory nor a fixed message. The compiler knows that a tile is currently local, that the next tile must arrive from another rank, and that holding fewer replicas saves capacity at the cost of more movement. It can therefore compare a synchronized all-gather against a sequence of sends and receives whose transfers overlap local work. The representation turns a collection choice into a memory-placement and time-ordering problem.

A shift can express a ring without naming one
Consider a weighted sum split across four ranks. A synchronous version can replicate one input and shard another, compute locally, then all-reduce the partial outputs. A shifted version shards the input that would otherwise be replicated and rotates its pieces. At each loop step, a rank consumes one local or newly arrived tile while sending another tile onward. Storage falls because no rank needs the entire input, and communication overlaps with arithmetic when tile timing permits.
CommIR expresses that schedule by shifting one local loop relative to a parallel loop. The compiler lowers the annotation into sends, receives, and local operations. Applying shifts to arbitrary loop dimensions can recreate a ring, head-parallel, or hybrid attention schedule without a separate handwritten implementation for each named algorithm. It can also produce a schedule that mixes patterns at the intra-node and inter-node levels, which matters because NVLink and RoCE differ by an order of magnitude in bandwidth and latency behavior.
The search space grows quickly. Mercury uses legality checks, memory constraints, and execution measurement to tune candidates. It compiles tensor operators written in a Python-embedded domain-specific language, lowers local work through PyTorch and TorchInductor components, and uses CUDA 12.6 with NCCL 2.26.2 in the reported implementation. It also reasons about the layout between adjacent operators so that an individually fast operator does not force an expensive reshard before the next one.
This last point is important for systems buyers. A kernel benchmark can reward a layout that makes the next layer slower. Mercury’s one-layer Llama experiments include QKV projection, attention, output projection, and MLP operators, and search resharding choices across them. The paper reports up to 1.62× improvement over a best 3D-parallel baseline for those model-level cases. The unit remains one Transformer layer, but the optimization boundary is wider than an isolated attention kernel.
Where the speedup comes from
The evaluation uses L4, A100, and H100 GPUs. The systems differ in memory capacity, HBM bandwidth, local connectivity, and inter-node bandwidth: L4 uses PCIe locally, while A100 and H100 use NVLink; the reported network configurations range from 50 to 100 Gbps between nodes. Node count ranges from one to four, with one to eight GPUs per node. Attention covers multi-head and grouped-query variants, batch sizes 1 and 16, and context lengths from 4K to two million tokens. Linear operators cover all-gather plus GEMM and GEMM plus reduce-scatter.
Mercury leads the chosen baselines in the paper’s operator sweep. On H100, one multi-head attention configuration at batch size 16 reaches roughly 4× the reference. On A100, one all-gather GEMM case reaches 1.9×. The advantage is not uniform. It narrows on simple one-level topologies where a handcrafted method already matches the hardware, and it also narrows beyond a million tokens when attention computation dominates communication. The strongest gains appear when local and remote links form a hierarchy and the compiler can choose different schedules for each level.
The topology experiment illustrates that point. With a fixed amount of work per GPU, adding nodes increases both capacity and remote communication. With a fixed 32K total context, adding GPUs reduces local work and can expose communication overhead. Mercury reports a 2.91× average speedup across the evaluated H100 and A100 layouts, but the gap is smaller for configurations that are entirely within one node or entirely across equivalent links. A compiler earns its value when the topology is irregular enough that one global collective is a poor fit.
Memory is the other axis. At two million tokens on eight H100 GPUs in a two-node arrangement, the baseline schedules in the reported comparison fail their memory constraint. Mercury finds a plan that shards KV state and outputs more aggressively. The plan may communicate more, but it converts an out-of-memory outcome into an executable one. This is the clearest example of why remote HBM is presented as a memory level: the objective is not only lower latency, but also a feasible placement under a capacity limit.

What has not yet been established
The paper does not report production deployment, cluster-wide scheduling, or multi-tenant interference. Its measurements use controlled GPU groups and a bounded set of operator shapes. Network congestion, topology changes, runtime jitter, failures, and concurrent jobs can make a schedule chosen from stable microbenchmarks less reliable. An operator compiler will need either robust plans or a retuning policy when the remote-memory path changes during service.
Compile and tuning cost also matters. Mercury measures candidate execution to find a good schedule, and the paper focuses on resulting performance rather than the full operational cost of maintaining profiles for a rapidly changing model catalog. A serving platform must decide when a configuration is common enough to justify tuning, how to cache schedules, and how to fall back safely when model shapes or hardware differ from the profile.
The model-level experiment benchmarks one Transformer layer because layers share a configuration. That is a reasonable way to isolate operator and resharding effects, but it does not include full-request behavior, pipeline bubbles, KV-cache lifetime across tokens, host overhead, admission control, or service-level tail latency. The four-node upper bound also leaves open how plans behave on rack and pod networks with more hierarchy and contention.
Lastly, remote memory is not interchangeable with local HBM. Mercury exposes the distinction precisely so the compiler can schedule around it. The phrase should not be read as a claim that another GPU’s memory has local latency or coherence. The system emits explicit communication through NCCL and related kernels. Its contribution is making those transfers part of a unified program, not removing the physical cost.
Remote HBM also changes capacity accounting
A compiler-generated schedule can make a kernel feasible, but a cluster operator still has to decide who owns the borrowed memory. Capacity admission cannot count only free bytes at launch. It must reserve the remote allocation, the communication buffers needed to move it, and enough link service to meet the compiled overlap assumption for the lifetime of the kernel. Otherwise two individually valid schedules can overcommit the same donor GPU or saturate the path they both expect to hide.
This creates a different failure boundary from ordinary local HBM allocation. A donor reset, a link reroute, or another tenant’s collective can invalidate a schedule even when the compute GPU remains healthy. The runtime needs a way to reject the chosen plan before execution, select a more conservative candidate, or fall back to local recomputation and smaller batches. A compiled fast path without a valid fallback would turn a performance optimization into an availability dependency.
Isolation is equally important. Remote memory contents, access permissions, zeroing, and lifetime must follow the same tenant boundary as local device memory. Scheduler telemetry should attribute borrowed byte-seconds and path occupancy to the consuming job rather than to the donor that happens to hold the allocation. Without that accounting, a topology-aware compiler can improve application latency while making cluster utilization and chargeback less intelligible.
The practical interface is therefore richer than a map of link bandwidths. It should expose reservable capacity, expected contention, failure domains, and the confidence interval around transfer performance. Mercury demonstrates how to search once those constraints are known. A production system must also refresh them and invalidate cached schedules when the placement or competing traffic changes. The useful abstraction is not unlimited remote HBM. It is a bounded, revocable memory level whose service contract the compiler can test.
The compiler needs a topology contract
The practical adoption question is what the runtime must promise to the compiler. A schedule chosen for eight H100s with a particular NVLink and RoCE layout can degrade after a placement change. The compiler therefore needs a topology descriptor that includes bandwidth, latency, congestion domain, memory capacity, and allowed communication primitives. The scheduler must either preserve that contract or trigger a new plan.
Operators should also expose an envelope of safe schedules rather than a single winner. One plan may minimize median latency, another may use less memory, and a third may tolerate network variance. Mercury’s design-space plot already contains a latency-memory Pareto frontier. A production system could add compile time, energy, and sensitivity to link slowdown, then choose according to service objectives.
We believe Mercury’s durable contribution is the representation. Multi-GPU optimization has been divided between tensor compilers that understand local loops and distributed libraries that understand collectives. CommIR places remote data ownership and movement in the loop program, so the compiler can search both sides of that boundary. The reported speedups show that the additional choices matter. The next proof is operational: preserve those gains when topology, traffic, and model shape change faster than engineers can retune them by hand.
Source and attribution
This article is an editorial summary prepared by Silicon & Systems. It restates the cited paper’s mechanisms, measurements, and limitations in our own words. No text, tables, or figures from the paper are reproduced; both figures and the card image were created for this article. The authoritative paper is available through ACM Digital Library under CC BY-NC-SA 4.0. Copyright (c) 2025 the authors.