A wafer-scale accelerator looks like a shortcut around distributed computing: keep the model on one piece of silicon and avoid a cluster network. The OSDI 2025 WaferLLM paper shows why that description is incomplete. Cerebras WSE-2 places 850,000 cores and 40 GB of aggregate SRAM on a two-dimensional mesh, but each core owns only 48 KB and communicates through a router with a limited number of configured paths. The device is physically one wafer and logically a very large distributed-memory machine. Software written for a GPU’s shared HBM can therefore run badly even when the wafer offers far more aggregate bandwidth[1].

WaferLLM starts from this mismatch, not from a new batching policy. Its PLMR model names four constraints: massive parallelism, non-uniform access latency, constrained local memory and limited routing resources. Prefill and decode are then rebuilt around those constraints. The result matters beyond one product because it asks a useful procurement question: when an accelerator changes the memory topology, how much of the advertised hardware advantage survives before the runtime is redesigned?

One package, distributed memory

The WSE-2 cited by the paper runs 850,000 compute cores at up to 1.1 GHz and provides 22 PB/s of aggregate on-chip memory bandwidth[1][2]. Those totals are striking, but an LLM kernel cannot consume aggregate bandwidth by issuing arbitrary loads. A datum placed on a distant core crosses more mesh hops than a local datum, and a route may require scarce router state. The paper reports a local-to-remote access-latency range that can differ by as much as 1,000×. It also notes that fewer than 25 hardware-assisted routes are available per core on the evaluated device.

This produces a different optimization target from a GPU. A GPU kernel normally tiles work to reuse data in registers and shared memory, then treats HBM as a large, uniformly addressable backing store. On the wafer, every tile has an address and a physical position. A placement that minimizes capacity pressure can lengthen communication paths; a schedule that exposes more parallelism can exceed route limits; replication can shorten distance but consume the 48 KB local budget. These are coupled constraints rather than independent knobs.

WaferLLM uses PLMR as a rejection test for familiar designs. Treating all local SRAM as shared memory fails because remote access cost grows with distance. Applying a conventional distributed algorithm also fails when it assumes thousands of participants with capable collective engines rather than hundreds of thousands of tiny routers. The useful abstraction is therefore not “one giant GPU” or “one giant cluster.” It is a mesh whose work and data must be placed together.

WaferLLM’s PLMR model rendered as an original editorial figure. The wafer-scale mesh is governed by four coupled constraints: massive core parallelism, up to a 1,000× access-latency gap, tens of kilobytes of local SRAM per core and fewer than 25 configured routes per core. The geometry is conceptual rather than a product photograph or manufacturing floorplan. Original figure created for this article.

Prefill as mesh-native matrix multiplication

Prefill is dominated by matrix-matrix multiplication. WaferLLM’s MeshGEMM breaks matrices into fine-grained blocks distributed over the two-dimensional core grid. It combines cyclic shifting with an interleaved communication order. Cyclic movement lets each core see the blocks required for a complete product while keeping its working set bounded. Interleaving avoids sending every long-distance message at the same moment, which shortens the critical path and reduces pressure on the limited route table.

The important point is not that the paper invents another GEMM. SUMMA and Cannon already distribute matrix multiplication, but their communication patterns were designed for different network and memory assumptions. WaferLLM evaluates whether each pattern respects the local-memory budget, route count and distance-dependent latency at wafer scale. MeshGEMM is the pattern that satisfies all three while expanding across hundreds of thousands of cores.

Tensor placement also removes a cost that is tolerable on GPUs but expensive on a mesh: transpose-like reshuffling between operators. WaferLLM keeps tensors in layouts that successive layers can consume directly. This is a systems result disguised as a kernel result. The winning schedule is the one that minimizes movement across the whole inference path, not the one with the most elegant isolated matrix primitive.

Decode needs replication, not only partitioning

Decode changes the arithmetic. Each step has a query length of one, so there is not enough independent work in the query dimension to occupy the wafer by partitioning alone. WaferLLM instead uses fine-grained replication and a mesh-specific matrix-vector multiply. Weights and partial work are distributed so that many cores participate without forcing every result through a single sequential reduction.

MeshGEMV uses a K-tree all-reduce. A pipeline reduction keeps route count low but takes a long sequential path; a ring has the same broad problem because partial sums traverse many participants. A balanced K-tree groups reductions into phases. Increasing K shortens depth but consumes more configured paths, so the implementation chooses K=2 for the evaluated WSE-2. That explicit trade is representative of the paper: no topology is universally optimal once router state is a resource.

KV-cache management follows the same logic. Concatenating new keys and values behind one contiguous region, as a GPU implementation might, skews capacity and work toward particular cores. WaferLLM shifts cache blocks across the mesh so that token growth remains balanced. The paper reports 360–385× more token capacity from this shift-based placement than the compared baseline layouts. That number is not extra physical memory; it is the difference between evenly usable local memory and capacity stranded behind an incompatible mapping.

What the measurements establish

The primary testbed is a Cerebras WSE-2. GPU comparisons use A100 accelerators made on the same 7 nm process, with up to eight GPUs connected by NVLink in one node and a second eight-GPU node reached through InfiniBand. SGLang is the GPU serving stack. The study evaluates complete LLaMA3-8B and LLaMA2-13B models; CodeLLaMA-34B and Qwen2-72B use representative layer subsets where a whole model does not fit on one WSE-2. The central metric is throughput per request, the reciprocal of time per output token, rather than aggregate batch throughput.

Against GPU-oriented or distributed-memory compilers mapped onto WSE-2, the gap is large. For short-output workloads, WaferLLM reports an average 160× advantage over T10 and 625× over Ladder. The comparison is useful chiefly as a software-abstraction experiment: both baselines leave most of the wafer idle because their memory model does not match the machine. It should not be read as evidence that one compiler revision made the silicon hundreds of times faster.

The A100 comparison is more relevant to serving decisions. Across the reported input/output lengths, WaferLLM produces 10–20× higher end-to-end per-request throughput than the best tested SGLang configurations and about 30–40× versus one A100. Energy efficiency is approximately 2–2.5× better at the complete-model level. MeshGEMV alone can look much more dramatic, reaching 606× lower latency and 16× better energy efficiency than one A100 in selected cases, but end-to-end inference also contains operations and software paths that do not scale by that factor.

Scaling is not perfectly monotonic. Enlarging the core grid adds communication distance, and some decode configurations slow when the all-reduce path grows faster than useful work. The paper’s best core count therefore changes by model and phase: for LLaMA3-8B, for example, the selected prefill grid is larger than the decode grid. A wafer is not a resource that should always be saturated. It is a topology whose efficient footprint depends on the operator.

The comparison that the paper does not make

WSE-2 and an A100 cluster differ in much more than process node. The wafer has far more silicon area and aggregate on-chip SRAM bandwidth, while the GPU system has off-package HBM capacity, mature compilers, continuous batching and a network that can scale capacity across machines. Throughput per request answers how quickly one request advances. It does not report tokens per rack, tokens per dollar, concurrent-user goodput under a latency SLO or failure-domain cost.

The paper is candid about software immaturity. WaferLLM contains roughly 7,000 lines of Cerebras CSL and 2,000 lines of Python, and some model operations do not yet exploit the wafer as effectively as its GEMM and GEMV kernels. Models larger than the 40 GB aggregate SRAM still require multi-wafer execution or layer-wise handling. The evaluation also predates the broad range of production GPU kernels and speculative-decoding methods available today. Any buying decision would need the same model, precision, request mix, batching policy, power boundary and availability target on both systems.

Yield and resilience require separate accounting as well. Wafer-scale systems route around defective regions, but a wafer remains a larger fault and service unit than one GPU. The paper reports that commercial wafer-scale devices can achieve high functional area through redundancy and remapping; it does not convert this into fleet availability or replacement economics. A low-latency request is valuable only if enough serving capacity remains during maintenance and faults.

A deployment checklist for topology-specific runtimes

The durable lesson is to make topology a first-class software input. A runtime for a new accelerator should expose local memory per compute element, access cost as a function of distance, route or link-state limits, supported collective shapes and the cost of changing placement between operators. Peak FLOPS and aggregate bandwidth cannot predict utilization without those fields.

Operators should then measure three layers separately. Kernel efficiency shows whether the local implementation is healthy. Per-request latency shows whether the topology accelerates the critical path. SLO-valid fleet throughput shows whether concurrency, failures and queueing preserve the advantage. WaferLLM is strongest on the first two. A production evaluation must add the third, along with model-loading time and the behavior of jobs that exceed on-wafer capacity.

Finally, compiler portability should be treated as a budget rather than an assumption. PLMR applies to other mesh-like devices only after its parameters change. A device with larger local SRAM can replicate more aggressively; richer routers can reduce K-tree restrictions; chip-to-chip meshes add a new latency tier. The concepts transfer, but the chosen schedule does not. The correct contract is that the compiler re-derives placement from measured topology instead of carrying a WSE-2 schedule to a different machine.

The system decision

WaferLLM does not prove that wafer-scale hardware replaces GPU clusters. It proves that a wafer cannot be evaluated with a GPU runtime and then dismissed for low utilization. Once data placement, route state and phase-specific parallelism match the mesh, the hardware delivers a substantial per-request latency and energy advantage on the tested models.

For infrastructure teams, this moves the question from chip specifications to software ownership. Buying a nonstandard memory topology means owning a compiler and runtime that understand it, including placement across model revisions. If a vendor supplies only peak numbers and a compatibility layer, much of the hardware can remain inaccessible. If the runtime exposes and optimizes the real topology, the wafer becomes more than a large die: it becomes a coherent inference system whose critical communication stays on silicon.

This article is an independent editorial digest of the OSDI 2025 paper[1]. The prose and figure were created anew for Silicon & Systems; no paper figure or table was reproduced. Factual claims and measurements are attributed to the authors’ published evaluation. Copyright in the original paper remains with its authors and the USENIX proceedings (2025).