Every LLM request begins with prefill and continues through decode, but the two phases stress a GPU differently. Prefill processes many prompt tokens in parallel and tends to consume arithmetic throughput. Decode produces one token per sequence and repeatedly reads weights and KV cache, so memory bandwidth dominates. Modern serving systems place both phases in one hybrid batch, yet their attention kernels can still run as separate specialists. One finishes with bandwidth idle; the other finishes with compute idle.

POD-Attention moves overlap inside the kernel. It builds on FlashAttention-style tiling, represents prefill and decode work as cooperative thread arrays, and schedules them with awareness of the streaming multiprocessor. The goal is not simply concurrent launch. It is to place complementary work on the same SM without letting one operation’s straggling CTA or shared-memory demand block the other.

Why two concurrent kernels do not guarantee overlap

CUDA streams can launch prefill and decode kernels concurrently, but hardware does not guarantee that their CTAs occupy every SM in a balanced mix. One kernel can fill the available SMs before the other is admitted. Horizontal fusion places operations in one launch but may assign a fixed resource partition that performs poorly as batch size and context length change.

Wave quantization creates another loss. A kernel whose CTA count slightly exceeds a multiple of available SM capacity needs an additional partial wave. A small increase in work can therefore produce a large latency step. The paper reports cases where stream concurrency gains up to 20% when it fills otherwise idle capacity, but it cannot reliably create SM-level co-location.

Intra-thread fusion guarantees co-location by making each thread handle both operations, but prefill and decode use different tile sizes and resources. The slow part can hold the whole fused unit, and tuning must cover a large space of batch and context combinations. POD-Attention chooses CTA-parallel fusion so GPU hardware can schedule completed units while software controls which operation each virtual CTA performs.

POD-Attention’s mechanism and evidence boundary. Compute-heavy prefill CTAs and bandwidth-heavy decode CTAs are co-located through an SM-aware fused kernel. The paper reports up to 59% faster attention with a 28% mean, and up to 22% end-to-end throughput improvement in the evaluated serving setup. Original figure created for this article.

Virtual CTAs and SM-aware balance

POD-Attention launches a regular grid but maps its CTA identifiers to virtual prefill and decode CTAs. A software scheduler inside the fused kernel chooses the next operation while respecting the resources available on each SM. This allows a completed decode unit to make room for prefill work, or the reverse, instead of waiting for an entire kernel wave.

Tile size is a real constraint. Larger decode tiles can raise memory-bandwidth utilization, reaching as high as 70% in the paper’s explored configuration, but consume more shared memory and can reduce occupancy. Prefill benefits from tensor-core-friendly tiles. The implementation hand-tunes shared-memory use and limits prefill splitting so one chunk does not create excessive CTAs that crowd out decode.

The fusion keeps the mathematical attention results unchanged. It changes the order and placement of independent tiles, not the model or approximation. Correct synchronization is required where a tile’s partial results meet, and the scheduler must avoid deadlock when operations have different progress. These implementation details are why a generic kernel-fusion pass is not enough.

Long context changes the payoff

For short contexts, linear layers can dominate an iteration and attention fusion affects only a small fraction of end-to-end time. As KV cache grows, decode attention reads more memory and prefill attention processes larger prompt chunks. The share of iteration time spent in attention rises, and complementary resource use becomes more valuable.

The evaluation includes Yi-6B, Llama-2-7B, and Llama-3-8B configurations on A100 GPUs, with context lengths extending from 4K to 32K in the online workloads. One trace contains 2,000 internal requests and another uses 2,000 arXiv summarization requests. Their prefill-to-decode ratios differ, giving the fused kernel distinct mixes rather than one fixed synthetic batch.

Against separate FlashAttention and FlashInfer paths, POD-Attention cuts attention time by as much as 59%, with a reported mean gain of 28%. Integrated into Sarathi-Serve, it improves end-to-end throughput by up to 22% while reducing time to first token, time between tokens, and request execution latency in the tested cases. The application gain is smaller because linear operations, scheduling, sampling, and other work remain.

The denominator behind 59%

The 59% result is attention-kernel speed, not whole-service throughput. It appears under a selected batch, context, and tile configuration. The 28% mean better represents the tested range, while the 22% maximum describes an end-to-end serving outcome. Using the kernel maximum as a capacity forecast would count time the service never spent in attention.

The baseline combines independently optimized kernels, so POD-Attention’s advantage includes better coordination, not a claim that FlashAttention or FlashInfer is poorly implemented. New kernel versions, different GPU architectures, or an inference engine with another batching policy can change the gap. H100 and later GPUs alter tensor-core throughput, memory bandwidth, shared-memory capacity, and scheduling behavior.

Latency improvements also depend on the queue. A faster hybrid iteration can reduce token gaps, but a scheduler that admits oversized prefill chunks may still create head-of-line blocking. Kernel and scheduler need to agree on chunk size and SLO priority. Optimizing either in isolation can move the bottleneck.

Resource complementarity as a scheduling signal

GPU serving schedulers commonly track token counts and memory capacity. POD-Attention suggests adding resource shape. Two batches with the same number of tokens can use compute and bandwidth differently depending on phase and context. Pairing complementary operations may finish sooner than grouping identical work, even before model-level parallelism changes.

The signal must remain simple enough for online use. Exhaustively profiling every possible batch is impractical. A system can bucket prefill chunks and decode groups by expected CTA count, shared memory, compute intensity, and bytes read. The kernel can expose a small set of supported mixes and their measured latency. The outer scheduler then chooses among validated operating points.

This is particularly useful for disaggregated serving. A dedicated prefill pool and decode pool cannot fuse work on the same GPU, so it trades this intra-device complementarity for independent scaling and KV transfer. The better architecture depends on network cost, load shape, context length, and whether each pool can stay occupied. POD-Attention strengthens the case for colocated hybrid batching where long contexts make attention dominant.

Isolation and predictability

Fusing two latency classes can improve average efficiency while complicating predictability. A decode request with a strict time-between-token objective now shares an SM with prefill work. The implementation must bound how much prefill enters one iteration and how long a fused CTA can run. Tail latency, not mean kernel speed, should drive the limit.

Admission control should also handle unsupported shapes. An extreme context, unusual head dimension, or quantized attention path may fall back to separate kernels. Capacity planning must include the fraction of traffic that takes the optimized path and the performance discontinuity at fallback.

Debugging becomes more specialized. GPU traces show one fused kernel rather than two recognizable launches. The runtime should report the virtual CTA mix, tile choices, achieved occupancy, memory throughput, and reasons for fallback. Without those counters, a scheduler change can look like an unexplained kernel regression.

A reproduction and procurement test

Teams should reproduce three layers. First, compare separate, stream-concurrent, and fused attention with fixed inputs while recording CTA waves, SM occupancy, tensor-core activity, memory bandwidth, and shared-memory use. Second, integrate the kernel into the actual serving scheduler and test request traces across 4K to 32K or the production context range. Third, report SLO-valid throughput and p99 token latency under load.

The test should vary prompt chunks, decode batch, model, precision, and GPU generation. It should record the percentage of iterations that use POD-Attention and the reason for every fallback. Power is relevant because higher simultaneous resource use can increase board demand even when requests finish sooner.

For procurement, the result argues against comparing accelerators only by peak FLOPS or HBM bandwidth. A balanced device and software stack can exploit compute and bandwidth at the same time. The value appears only if the kernel and scheduler expose enough flexibility to construct complementary work.

The system decision

Hybrid batching does not automatically create hybrid execution. Separate prefill and decode kernels can leave complementary GPU resources idle even when requests share one batch. POD-Attention makes the overlap explicit at CTA and SM granularity, producing a meaningful attention speedup and a smaller but real service-level gain in long-context tests.

The operational decision is whether attention occupies enough of the request path, and whether the service can bound the fused iteration for tail latency. Where long contexts and mixed phases dominate, co-locating compute-heavy and bandwidth-heavy tiles can raise capacity without changing the model. Where linear layers or strict isolation dominate, the additional kernel complexity may not pay. The required evidence is an end-to-end SLO curve, not the fastest attention microbenchmark.

This article is an independent editorial digest of the ASPLOS 2025 paper[1]. The prose and figure were created anew for Silicon & Systems; no paper figure or table was reproduced. Results retain the paper’s models, GPUs, traces, context lengths, and baselines. Copyright in the original paper is held by its authors and publication rights are licensed to ACM (2025).