A GPU does not need to be full for memory to be the bottleneck. Graph features, recommendation embeddings, and key-value (KV) cache blocks can remain in host memory because their complete working sets are larger than device memory. Each batch then brings only the required tensors across PCIe. Compute and transfer can overlap, but overlap stops hiding latency when the next tensor arrives after the current kernel has finished.

Compression appears to create bandwidth, although the usual version changes the data. Quantization reduces the number of bits per value and may be acceptable when accuracy has been characterized for a particular model. It is a harder operational choice when one serving platform runs many models, input distributions drift, or a correctness requirement permits no numerical change. Invariant Bit Packing (IBP) asks whether exact values can travel in fewer bits because some bit positions are constant across a group.[1]

The bottleneck sits between two fast memories

GPU memory bandwidth and arithmetic throughput have advanced faster than the host-to-device path. A workload that fetches data on demand therefore encounters three distinct limits: how much of the working set fits in GPU memory, how much useful data PCIe delivers per second, and how long reconstruction occupies the GPU. Caching improves the first limit. Prefetching hides part of the second. Neither removes bytes from a compulsory miss.

The paper studies three workloads that expose different forms of this problem. GNN training gathers feature vectors for sampled vertices. A deep-learning recommendation model looks up sparse embedding rows. Long-context LLM inference can move KV-cache tensors that are not kept locally. Their access patterns differ, but all transfer structured numerical arrays whose bit patterns may contain redundancy.

Lossy methods trade representation error for a predictable reduction in bytes. That trade is not uniform across these workloads. A quantization format that preserves one graph or model may change another model’s ranking or generated distribution. A platform team must then validate accuracy per workload and repeat that validation after model changes. IBP removes this accuracy branch by reconstructing the original bit strings exactly.

Packing invariant positions rather than coding every value

IBP groups values and examines each bit position across the group. Positions that have the same value everywhere are recorded once as metadata. The remaining positions are packed into a dense payload. The host holds that packed representation; a small description of invariant positions is available to the GPU. On a miss, the payload crosses PCIe and a GPU kernel inserts the omitted bits before the consuming operator uses the tensor.

IBP separates shared bit positions from a compact payload, transfers fewer bytes over PCIe, and reconstructs the exact tensor with warp-parallel operations. Original figure created for this article.

This scheme differs from a general-purpose codec in where it spends work. Dictionary construction, variable-length parsing, and large temporary buffers can erase the transfer savings when tensors are needed immediately. IBP favors fixed bit operations and a layout that many GPU lanes can unpack in parallel. Its metadata is on the order of kilobytes rather than a second large compressed-state structure. The decompressor sits just before use, so the system does not first expand the entire working set into scarce GPU memory.

The invariant mask is also a compact statement about the data. Sign and exponent fields can repeat across floating-point values with similar scale, while integer identifiers can share high-order zeros. The amount of removable data therefore depends on representation, grouping, and distribution. IBP is not a constant-ratio link codec. It is a workload-aware representation placed on the host-to-device path.

Zero copy matters as much as compression ratio

A useful compression ratio is insufficient if the CPU must copy every request into a staging buffer. The paper arranges compressed data so the GPU can read it through asynchronous transfers and reconstruct it without putting the host processor in the critical path. This matters for sparse workloads, where many small gathers can otherwise turn allocation and copying overhead into the new bottleneck.

The same accounting applies on the GPU. Decompression consumes memory bandwidth, instructions, and scheduling slots. It wins when the PCIe time removed exceeds those costs and when the reconstructed output can flow directly into the next operator. It can lose on data that compresses poorly, on tensors already resident in HBM, or on kernels whose compute time fully covers transfer.

Thus the deployment decision needs three measurements for each tensor class: bytes saved, transfer time avoided, and decode time added. A single corpus-wide compression ratio conceals the tensors that do not benefit. A practical implementation should sample inputs, select IBP only above a measured break-even point, and retain an uncompressed path.

Three workloads, three meanings of speedup

On an NVIDIA A100, average speedups are 74% for GNN training, 180% for DLRM embedding lookup, and 25% for LLM inference.[1] These numbers are intentionally not interchangeable. The GNN result includes a training pipeline in which feature fetching and computation interact. The DLRM result measures an embedding lookup path with substantial sparse movement. The LLM result depends on the fraction of inference time spent retrieving KV data rather than executing attention and projection kernels.

The 180% wording means the measured lookup becomes 2.8 times the original throughput or completes in roughly 36% of the time if work is otherwise equivalent. It does not mean an end-to-end recommendation service becomes 2.8 times faster. Likewise, a 25% LLM inference gain can be material for a memory-tiered long-context service, but it should not be applied to a configuration whose KV cache already fits in GPU memory.

The A100 platform also fixes the balance among PCIe bandwidth, HBM bandwidth, and decompression throughput. A newer accelerator, CXL-attached memory, a different CPU socket topology, or NVLink-connected memory tier changes that balance. IBP’s mechanism can remain valid while its break-even tensor size and compression threshold move.

Exact reconstruction narrows, but does not remove, risk

Losslessness simplifies model governance because the decompressed tensor matches the source bits. It does not automatically prove system correctness. The implementation must cover every data type, alignment, tail group, and metadata boundary. A corrupted invariant mask can affect many values at once. Checksums and fallback decoding should therefore be part of the storage format, not added only after an incident.

Sampling introduces another operational question. A mask learned from one region of a dataset may not compress later regions equally well. The paper demonstrates streaming support through sampling, but production telemetry should still track realized ratio and decode cost over time. A workload should leave the compressed path when the ratio falls below its break-even threshold.

Memory capacity must also be counted end to end. The compressed host copy, GPU metadata, in-flight payloads, and reconstructed output can coexist briefly. Queueing several transfers to saturate PCIe may increase that transient footprint. Capacity planning that credits only the compressed payload can overcommit pinned host memory or GPU scratch space.

What to measure before adopting IBP

A representative test should separate compulsory transfers from cache hits. Report the tensor classes, data types, group sizes, compression-ratio distribution, and the share of requests that use the compressed path. Compare against the same cache and prefetch policy without compression. Otherwise, a better placement policy can be mistaken for a codec gain.

Latency distributions matter alongside throughput. Just-in-time decode can add a small fixed cost to every fetched tensor even when average throughput rises. Interactive LLM serving should expose time to first token and inter-token latency by context length. GNN and recommendation workloads should report batch-tail latency, not only epoch or lookup averages.

Power and contention complete the evaluation. Decompression uses GPU instructions that could overlap with model kernels, while a smaller PCIe payload can reduce I/O energy and release link capacity for neighbors. Multi-tenant tests should show whether one compressed workload improves the node or merely shifts interference from PCIe to SMs and HBM.

The system decision

IBP makes a useful separation: numerical precision and transfer size do not have to be the same decision. When data contains stable bit positions, the system can preserve the model’s exact representation and still send fewer bytes. That is especially attractive for shared platforms that cannot revalidate an accuracy trade for every model revision.

The method does not make PCIe disappear. It adds a codec whose value depends on the tensor, accelerator, and placement policy. The right control loop measures those conditions and selects compressed or direct transfer per tensor class. For buyers and operators, the headline gains should therefore be translated into a break-even map: which working sets spill, which values actually compress, and which GPU kernels have enough idle time to reconstruct them. That map determines whether lossless packing creates useful bandwidth or only another stage in the pipeline.

This article is an independent editorial digest of the EuroSys 2026 paper[1]. The prose and figure were created anew for Silicon & Systems; no paper figure or table was reproduced. Reported results preserve the authors’ A100 platform, workloads, and baseline definitions. The original paper is distributed under CC BY 4.0; copyright is held by the authors (2026).