Deploying an LLM is not one sizing exercise. For every workload, an operator must select a model, accelerator, precision, batching policy, and runtime. The remaining choice spans five parallel dimensions: tensor, pipeline, expert, context, and data. The selection must fit memory, match the physical fabric, satisfy time-to-first-token and incremental-token limits, and minimize cost. A reasonable choice for interactive chat can be wasteful for offline scoring. A configuration that wins for prefill can be wrong for decode. Changing the model from dense to mixture-of-experts can reverse the best cross-host strategy.
Meta’s MLSys 2026 industry paper describes how its inference team handles this combinatorial problem for Llama services reaching nearly one billion monthly active users[1]. The team builds a lightweight simulator from measured operator performance, assembles end-to-end phase models, eliminates configurations that violate memory or latency constraints, and ranks the remainder by cluster throughput or cost. A typical exploration crosses thousands of runtime and hardware options with roughly a thousand parallel configurations, bringing the total into the millions.
The headline should not be that one runtime or accelerator wins. The paper’s five findings are conditional. Phase separation saves capacity for latency-sensitive online services but adds little to offline throughput. Different accelerator types can lower projected total cost when prefill and decode favor different resources. MoE scales across ordinary networked hosts only when expert parallelism replaces expensive cross-host tensor collectives. Scale-up fabrics matter most when large dense models need wide, latency-sensitive parallelism. The durable contribution is the method that decides which condition applies.
A deployment is a constrained optimization problem
The simulator maximizes cluster queries per second while keeping TTFT and time-to-incremental-token below service limits. Its inputs include model architecture, request lengths and rates, hardware, parallel strategy, runtime design, and SLOs. Even before model compression or speculative decoding, those variables create a large discrete space. An online 70B service with a 2K-token prompt, 150-token response, and 8K context behaves differently from an offline job with a 1K-token response and no interactive deadline.
Meta starts with more than 100,000 microbenchmark records per hardware platform. The database covers GEMM, attention, all-reduce, all-to-all, and other common operators across shapes. Piecewise interpolation estimates an operator when an exact point is absent. A model assembler then maps model and parallelism choices into operator shapes, follows the critical path, and includes communication and runtime overhead. Separate prefill and decode estimates feed either a continuous-batching model or a disaggregated model.
Measured and predicted latency generally differ by no more than about 5% in the reported validation. The simulator explores a design in minutes, which makes repeated evaluation possible as workloads and hardware change. That accuracy mainly covers median and mean performance around measured points. The authors explicitly note that network jitter can affect p99 behavior, so the model is a configuration selector rather than a substitute for production canaries.
Search pruning is part of the system. Plans that exceed memory or miss the SLO are removed early, and common eight-GPU topologies favor power-of-two parallel degrees. The system can model nonstandard pod sizes, speculative decoding acceptance rates, empirical MoE routing, power-capped hardware, and total cost. This structure turns deployment from a collection of expert heuristics into a reproducible experiment with stated constraints.

Separate prefill and decode only when the SLO pays for it
Prefill processes many prompt tokens in parallel and is often compute-bound. Decode repeatedly reads weights and KV state for relatively little arithmetic, making memory bandwidth more important. Continuous batching lets both phases share a pool and is operationally simple. Disaggregation gives each phase its own hardware count, batch size, and parallel plan, but introduces rate matching, KV transfer, routing, and another capacity boundary.
For the paper’s strict online scenarios, disaggregation produces 1.51.8× the throughput of continuous batching for Llama 3 70B and 1.82.2× for 405B while meeting the same latency limits. Meta reports moving most online services to a disaggregated runtime and obtaining about 30% capacity savings. The gain comes partly from allowing decode batches much larger than a shared runtime could accept without delaying prefill.
Offline generation gives a different answer. Without a tight token-latency target, both runtimes converge toward deep pipeline parallelism and large batches. In the evaluated 70B case on one hardware type, continuous batching slightly exceeds disaggregation. The extra operational surface of a split runtime is then difficult to justify. Runtime architecture should therefore be selected after the latency constraint is fixed, not adopted as a general mark of sophistication.
Parallelism also spends latency headroom. For online prefill on one platform, a TP4-PP2 plan lowers latency but delivers 20% less throughput than TP2-PP4 because more all-reduce traffic consumes the gain. The best plan does not minimize latency. It uses the available TTFT budget to reduce communication and process more requests. Decode may choose TP8 on the same service to aggregate HBM bandwidth. Phase separation is valuable because it permits these two plans to coexist.
Hardware heterogeneity is a ratio, not a shopping list
Disaggregation creates a natural opening for heterogeneous accelerators. A compute-rich device can run prefill while a device with more HBM bandwidth runs decode. In the paper’s 405B online example, the prefill-favored platform delivers 88 phase QPS but only 77 for decode, while two bandwidth-rich platforms reach 276 or 277 decode QPS. Combining the first type for prefill with the second for decode matches the best homogeneous end-to-end throughput.
Meta’s cost model projects a 15% to 25% total-cost improvement for such mixes under its actual price structure. This is a modeled opportunity, not a universal hardware discount. The required prefill-to-decode ratio changes from 0.88 for one homogeneous pair to more than 3.1 when hardware types are mixed. Traffic shape, token lengths, price, availability, power, software maturity, and transfer overhead can all change the result.
The fleet consequence is more demanding than assigning phases to labels. Capacity planning must maintain the correct ratio between two pools across geographically distributed datacenters. A shortage in either phase limits end-to-end throughput, and spare devices in the other pool do not repair it. Heterogeneity saves money only when the scheduler, model runtime, and supply plan preserve the phase ratio at the service’s arrival pattern.

MoE changes which communication is tolerable
A sparse MoE model activates only some feed-forward experts for each token. That can reduce computation, especially in prefill, but total parameter capacity and routing traffic remain. The paper compares a hypothetical 405B-scale MoE construction with a dense model to isolate system effects, then validates communication behavior on Llama 4 Maverick. The hypothetical model is not presented as a quality-equivalent replacement, so its raw throughput numbers should not be treated as a model recommendation.
Within one tightly connected host, tensor parallelism can work for dense and MoE models. Across hosts, its repeated all-reduce becomes expensive. Expert parallelism places experts on different devices and uses all-to-all to route tokens. In the Llama 4 Maverick decode experiment, extending TP8 to TP16 across two hosts lowers cluster QPS from 618 to 329. A hybrid plan that retains TP8 within a host and uses two-way expert/data parallelism across hosts preserves 608 QPS, only 2% below the single-host reference. Another four-way hybrid reaches 899 QPS, 45% above the TP8 reference by processing more tokens with less costly communication.
Routing balance remains a first-order input. In one production workload, the busiest of 128 experts receives about 7% of routed tokens. Prefill latency follows the most-loaded expert domain, so hot and cold experts should be placed together. During decode, cost depends strongly on how many distinct experts become active because each active expert can require another weight read. Three plausible routing models differ by as much as 3.5× in predicted cost at moderate batch sizes. A simulator that assumes perfect balance can select the wrong deployment.
Scale-up is necessary only after model and SLO force it
The paper compares conventional eight-card scale-out hosts with a projected 64-card scale-up pod for hypothetical 1.8-trillion-parameter dense and MoE models. Dense models that require wide tensor parallelism suffer on the ordinary inter-host fabric because all-reduce cost grows faster than useful work. The high-bandwidth scale-up fabric keeps wider TP feasible, although returns still diminish at 32 and 64 cards.
MoE has another path. Expert parallelism can scale decode across ordinary hosts because all-to-all routing is less demanding than the dense model’s wide all-reduce in the studied configurations. Scale-up still provides lower latency and stronger throughput, but scale-out remains viable and preserves smaller failure domains. The decision is therefore not “scale-up for every large model.” It is scale-up when model capacity or latency forces communication that the scale-out fabric cannot absorb.
Most numbers in this section are projections from the benchmark-driven simulator, including the next-generation platform and hypothetical 1.8T models. They are useful for architecture planning but are not equivalent to a production measurement. Hardware procurement should rerun the model with vendor-specific topology, power, price, availability, and workload traces, then validate the chosen frontier on representative systems.
A selected configuration has an expiration date
The optimizer produces an answer for a workload distribution, price table, hardware inventory, and software release. None is stationary. Prompt and output lengths drift after a product change, a quantized kernel changes the balance between compute and memory, and a temporarily scarce accelerator can reverse the cost ranking. The selected deployment should therefore carry the assumptions under which it won and a time at which those assumptions will be measured again.
Prediction error should enter that decision directly. A configuration that barely satisfies the SLO at the model’s median estimate is not equivalent to one that survives the upper error bound. Meta reports typical median and mean latency error within about 5%, but tail behavior and a newly introduced model can be less certain. The search can reserve a safety margin proportional to validation error, then reduce it only after canary traffic confirms the prediction on the actual serving stack.
Changing configurations also has a cost that static TCO misses. Moving weights, warming caches, rebuilding parallel groups, and draining requests consume capacity and can expose a transient latency tail. A controller should switch only when the expected saving over a defined horizon exceeds that transition cost. For heterogeneous deployments, it must also account for failure and maintenance: losing the less common accelerator type can strand the abundant tier even when total fleet capacity appears sufficient.
These requirements turn a one-time notebook result into a repeatable operating loop. Keep the workload trace and benchmark version with every recommendation, replay the incumbent as well as new candidates, validate the predicted SLO with shadow or canary traffic, and retain a known-safe rollback configuration. The important output is not merely the lowest-cost point. It is a decision whose evidence, uncertainty, switching cost, and expiry can be inspected by the team that operates it.
The operating artifact is the search process
Meta’s paper offers direction rather than a portable configuration table. Smaller models need fewer devices, so parallelism and scale-up matter less. A different request-length distribution changes the prefill-to-decode ratio. A different SLO changes the batch size that remains legal. Hardware prices can reverse a heterogeneous placement even when throughput is unchanged. Production tail latency can invalidate a plan whose median estimate looks correct.
An operator adopting this method needs versioned inputs: workload traces, SLOs, model graph, kernel and collective measurements, hardware topology, failure assumptions, power, and cost. Every runtime, compiler, firmware, and model change should identify which profiles are stale. The selected plan then moves through a canary that measures p50 and p99 phase latency, queueing, KV transfer, throughput, quality, and failure recovery.
We read the paper as an argument against architecture by slogan. Disaggregation, heterogeneous accelerators, MoE, expert parallelism, and scale-up fabrics are not independent upgrades. Each changes the feasible region of the others. A measured search process keeps those interactions visible and preserves the denominator behind every capacity claim. The best deployment is not the fastest component. It is the least costly configuration that continues to satisfy the service contract under the workload actually arriving.
Source and attribution
This article is an editorial summary prepared by Silicon & Systems. It restates the paper’s method, case studies, and limitations in our own words. No source sentences, tables, or figures are reproduced; both figures and the card image were created for this article. The paper is available in the MLSys 2026 proceedings. Copyright (c) 2026 the authors.