The decode phase of an LLM presents a deceptive utilization problem. Each request generates one token at a time, so the query dimension is only one element wide while the KV context can span hundreds of thousands of tokens. Attention must stream through that long context, yet a fixed partition can launch too few independent thread blocks to fill every streaming multiprocessor. The kernel is bandwidth-bound, but parts of the GPU are still idle.
Microsoft’s LeanAttention paper at MLSys 2025 focuses on this last, partially occupied execution wave[1]. It divides the context into hardware-sized units called LeanTiles and assigns a continuous stream of those units across the available compute resources. Unequal pieces can be combined exactly because online-softmax rescaling is associative. The result preserves mathematical attention while changing how work is distributed.
Prefill and decode are different shapes
FlashAttention earns its speed by tiling attention so intermediate score and probability matrices remain on chip instead of being written to global memory[2]. FlashAttention-2 adds parallelism over query length. This works well in prefill because a prompt supplies many query rows. In decode, however, the current query is a single token. Parallelism is limited mostly to batch and attention heads, neither of which can grow freely when KV caches already consume most of memory.
FlashDecoding addresses the problem by splitting the key/value context and reducing partial attention outputs. Its split count is fixed for a kernel launch. That creates quantization loss in the scheduling sense: the number of cooperative thread arrays does not divide cleanly by the number of SMs. Some execution waves are full; the last is only partially occupied. Increasing the split count fills more SMs but creates more partial outputs, consumes registers and adds reduction work. The best split therefore changes with context length, batch, head count and GPU model.
Heterogeneous batches make the mismatch worse. Continuous batching commonly combines requests with different context lengths. A fixed amount of work per split leaves some blocks with a long context segment and others with a short one. The kernel completes at the speed of the longest block, so average utilization can look reasonable while tail work still determines latency.
Exact softmax can be reduced in pieces
The obstacle to arbitrary partitioning is softmax. Each output depends on the maximum and exponential sum across the whole context, which seems to require equal, predetermined chunks. Online softmax already updates these statistics incrementally. LeanAttention observes that the rescaling operation used to merge partial outputs is associative: separately computed context ranges can be combined in any grouping as long as each carries its local maximum, normalization sum and weighted value.
That property lets the scheduler split one attention head into unequal work volumes. Each LeanTile is the smallest block that uses the GPU efficiently. A global work index maps tiles from batches, heads and context ranges onto CTAs, much like stream-K distributes matrix multiplication along its reduction dimension. CTAs pull the next tile rather than being pinned to a fixed split. When one request has more context than another, it contributes more tiles but does not leave a dedicated block waiting.
The final reduction rescales partial values to a common maximum and combines them. No token is discarded and no softmax term is approximated. This distinguishes LeanAttention from methods that change the attention function to gain speed. Exactness does not make it universally free: more tiles still mean more reduction state, and the chosen LeanTile size must match register, shared-memory and occupancy limits.

The hardware-aware part
LeanAttention derives tile count from the total work and the number of available compute units. It aims for near-complete SM occupancy, but occupancy here means that blocks are resident and useful work is distributed, not that tensor cores operate at peak arithmetic throughput. Decode attention remains dominated by memory movement for many shapes. Filling the machine reduces idle gaps; it does not repeal the memory-bandwidth limit.
The implementation also chooses between persistent and nonpersistent scheduling behavior. A persistent kernel can keep CTAs resident and pull work dynamically, which is useful for irregular batches, but holding resources too long can interfere with other kernels in a serving pipeline. The paper’s mechanism is therefore best understood as a scheduling primitive that a complete runtime should coordinate with batching and model execution.
Multi-GPU tensor parallelism changes the number of heads per device. Adding GPUs can reduce local work until each device has too few heads to fill its SMs, exactly where context-parallel LeanTiles are useful. The same design can run on NVIDIA A100 and AMD MI250X because it reasons about compute-unit count and work granularity instead of baking in one warp count. Portability still requires backend-specific kernels and tuning.
What the evaluation measures
The paper evaluates attention kernels across context length, batch size, head count and head dimension on A100 and MI250X accelerators. Baselines include FlashAttention-2, FlashDecoding and FlashInfer-style fixed-split behavior. Model-level experiments integrate LeanAttention into ONNX Runtime and use Phi-family configurations to measure decode and end-to-end inference.
Across the reported decode-attention tests, LeanAttention averages 1.73× lower latency than FlashDecoding. At 256K context length the gain reaches 2.18×. Some individual shapes improve more: the paper reports more than 3× in selected large-context, low-head-count cases where fixed splitting leaves many SMs unused. As batch size or head count grows, the baseline fills the GPU more naturally and the gap narrows.
The context trend is not simply “longer is always faster.” Longer context supplies enough tiles to occupy the machine, but it also increases memory traffic and reduction work. LeanAttention improves the slope relative to fixed splitting by distributing that work evenly. The absolute latency still grows with context length.
Heterogeneous batches provide a second source of gain. When the ratio between average and maximum context length falls, fixed partitions become increasingly imbalanced. LeanAttention’s global tile stream lets shorter and longer requests share compute units, so its relative advantage rises. This result connects the kernel to real serving, where equal-length batches are the exception rather than the rule.
Kernel speed is not service goodput
The headline number covers attention execution, not an entire serving system. A transformer step also runs QKV projections, output projection, feed-forward layers, normalization, sampling and framework code. The paper reports that decode attention can account for 50–60% of inference time in long-context configurations after other matrix multiplications are optimized. If attention is a smaller fraction in a particular model or batch, Amdahl’s law limits the end-to-end gain.
The model experiments confirm this dilution. End-to-end speedup varies with prompt-to-output ratio, generated-token count and the share of time spent in decode. Prefill-heavy requests obtain less value because LeanAttention targets query-length-one execution. A service whose users submit long documents but request one short answer may spend most time in prefill. A multi-turn agent generating thousands of tokens against a growing context is a better match.
SLO-valid throughput adds further constraints. Lower kernel latency can allow larger batches, but KV-cache capacity may cap batch size first. Dynamic batching, prefix reuse and prefill/decode disaggregation can alter which phase controls queueing. Operators must measure time between tokens and completed requests under the same arrival trace, not multiply a microbenchmark speedup by fleet size.
Exactness and numerical qualification
LeanAttention keeps the exact attention equation at the algorithmic level. In floating-point execution, changing reduction order can still change the last bits because addition is not associative numerically. This is normal for parallel reductions, but model qualification should compare output tolerances across precisions and long contexts. “Exact” here means no deliberate sparsification or softmax approximation, not bit-for-bit identity for every schedule.
The reduction state also needs careful bounds. Online softmax avoids overflow by tracking maxima, yet very long sequences and low-precision accumulation can expose numerical differences. A production kernel should test adversarial score ranges, masked tokens, grouped-query layouts and mixed context lengths. Correctness tests should accompany every architecture-specific optimization.
Energy claims need the same scope discipline. Better occupancy can finish a kernel sooner, while fully active SMs can also draw more instantaneous power. The paper’s appendix reports lower attention-kernel energy relative to fixed-split baselines as contexts grow. Fleet decisions should include HBM energy, idle power between tokens and the effect of serving more concurrent requests, measured at the node boundary.
A useful control policy
LeanAttention should not replace every decode kernel unconditionally. A runtime can select among FlashAttention-style, fixed-split and stream-K-style schedules from the actual shape. Small contexts or large batches may already fill the GPU and avoid the extra reduction. Long contexts, few local heads and heterogeneous batches are the high-value region.
The selector needs observable inputs: context-length distribution per batch, heads per device after tensor parallelism, head dimension, SM count, available shared memory and current GPU contention. It should benchmark more than average latency. P99 time between tokens, energy per generated token and interference with adjacent kernels reveal whether a persistent tile stream improves the service rather than only the isolated operator.
This policy also belongs near admission control. A request with a 256K context consumes both attention time and KV capacity. Faster attention may admit more such requests, but memory fragmentation or prefill queueing can still break the SLO. The scheduler should price the full request before deciding that a faster decode kernel created capacity.
The system decision
LeanAttention identifies a narrow but important source of waste: the last incomplete GPU wave in decode attention. Its contribution is to remove that waste without changing model semantics, using an algebraic property of online softmax to create a hardware-sized stream of work. The reported average 1.73× attention speedup is credible because it varies with the shapes that predict occupancy.
The broader lesson is that “memory-bound” is not a sufficient diagnosis. A kernel can be limited by memory bandwidth and still schedule that bandwidth inefficiently across compute units. Infrastructure teams should inspect both bytes moved and the distribution of work. When long contexts, low local head counts and irregular batches coincide, balanced exact tiles can reclaim capacity. When other phases dominate, the right answer may be a different runtime change entirely.
Source and copyright note
This article is an independent editorial digest of the MLSys 2025 paper[1]. The prose and figure were created anew for Silicon & Systems; no paper figure or table was reproduced. Measurements are reported with the scope used by the authors. Copyright in the original paper remains with its authors (2025).