A new high-speed network does not replace every old endpoint on the day it enters service. Storage servers, inference workers, and specialized machines can retain different NIC generations and transport capabilities. If one endpoint cannot produce packets accepted by the other’s RDMA hardware, a fast link alone does not preserve the intended communication path.

BURST, developed by ByteDance and Hunan University, addresses that asymmetric case[1]. One endpoint uses an Ethernet NIC without RoCE support, while the peer has a commercial RNIC. Instead of forcing both sides to abandon their RDMA-facing application interfaces, BURST supplies the missing transport behavior in a userspace process.

The important result is not that software has made all NICs equivalent. The paper shows where CPU execution, shared memory, DMA support, and queue parallelism can recover useful performance. It also contains a consequential distinction: a roughly fourfold improvement in measured KV-cache transfer does not mean a fourfold improvement in the complete inference response.

Ethernet connectivity is not RDMA compatibility

Ethernet describes the network carrying the traffic, but it does not guarantee that an endpoint implements RoCE transport semantics. An RNIC expects appropriate headers, sequence handling, memory permissions, and completion behavior. Sending ordinary traffic over the same physical link does not make it a valid remote-memory operation.

BURST targets communication from a non-RNIC to an RNIC, which the paper calls NR2R. This differs from optimizing transport between two RNICs or building a proprietary software transport between two ordinary NICs. The remote commercial device remains a participant with hardware-defined expectations, so the software endpoint must speak a compatible protocol.

That is particularly relevant when a disaggregated application crosses hardware pools. A prefill worker and a decode worker may not have identical networking capabilities. Similarly, replacing a portion of a storage fleet can leave frequent communication between new and retained machines. A compatibility layer can reduce the cost of those mixed paths without requiring an immediate fleet-wide replacement.

The correct alternative depends on the application. If an existing TCP implementation already meets the service objective with acceptable CPU use, a new transport introduces deployment and operational work. If fallback networking consumes excessive host resources or delays state transfer, preserving the RDMA path can have a more direct benefit. BURST provides evidence for the latter case under explicit test conditions.

A separate process without a separate application API

Applications continue to use the standard RDMA library interfaces, while a BURST provider and supporting driver connect those calls to the service process. Queue pairs, completion queues, and registered memory are shared or mapped so the service can process application work. Compatibility at the API level therefore does not mean that deployment requires no software installation or configuration.

The kernel remains involved in privileged resource registration and mapping. After those resources are established, the high-frequency data path operates in userspace through DPDK. Describing the system as kernel bypass should not be read as eliminating the kernel from memory protection, device setup, or every control operation.

The independent process can serve several applications and assign their queue pairs to worker threads. This makes resource sharing explicit rather than embedding a separate transport instance in every application. It also makes the service’s scheduling, failure behavior, and resource limits part of the host’s operational responsibilities.

Moving queue context into host memory avoids depending entirely on a NIC’s limited context storage. However, host memory is not costless or uniformly fast. Cache locality, NUMA placement, and access contention still affect how quickly software can find and update connection state. The design exchanges one set of hardware constraints for a software-managed resource budget.

Doorbells are a concurrency and recovery problem

An application must notify the service when it posts new work. A hardware doorbell produces a device-visible event; a software service needs an efficient way to discover that event across processes. Polling a separate notification structure for every queue pair becomes expensive as the population grows, while frequent kernel notifications reintroduce overhead.

BURST groups producers into shared doorbell queues associated with worker threads. Several application queues can feed one worker, while multiple workers provide parallel processing. This avoids making one global lock or one polling loop responsible for every active connection.

The paper describes a failure that is easy to miss in a throughput-only test. A producer can reserve a queue slot and then crash before committing it. If the consumer waits forever for that reservation to become valid, one failed application can stop unrelated work sharing the worker. The queue must preserve progress in addition to avoiding ordinary contention.

A timeout-based reclamation mechanism invalidates abandoned reservations so processing can continue. This should motivate failure-injection tests, not just a successful steady-state benchmark. An operator needs to verify application termination, slow producers, stale mappings, and service restart behavior under the actual process model before relying on a shared transport for multiple tenants.

Removing a copy and moving a copy are different changes

On transmit, BURST maps registered application memory so the NIC can obtain the payload without first copying it into a merged protocol buffer. Headers and payloads can be supplied separately for transmission. The software still prepares protocol information and tracks work completion; zero-copy payload handling does not mean zero CPU execution.

On receive, incoming packet buffers must deliver data into the application’s registered destination. BURST retains that copy. Intel DSA can perform it asynchronously and reduce CPU copying work, but the bytes still move. Calling the complete receive path zero-copy would obscure the operation that determines when the application can safely observe its data.

The completion notification must follow the required copy completion. An early completion would report success before the destination held valid contents. The implementation also restricts receive work requests to one scatter/gather entry for this path, a concrete compatibility limit that matters to applications using more complex buffer layouts.

The DSA evaluation uses a separate 100G testbed. It should not be merged with the 400G multi-NIC measurements as if all numbers describe one identical configuration. The benefit also depends on whether receive copying is the limiting stage; accelerating it has little effect when sending or packet processing already determines throughput.

Transmit and receive have different data-movement responsibilities. The transmit path can reference registered payload memory directly, while the receive path retains a buffer-to-application copy and must wait before reporting completion. DSA offloads the latter copy rather than eliminating it. Original figure created for this article.

GPU access and transport policy still need validation

The GPU transmit path uses peer-memory support to obtain the mappings needed for device DMA. The NIC then reads payload from GPU memory and headers from host memory. This is not evidence that every Ethernet NIC, GPU, driver, and PCIe topology supports the same access path without additional requirements.

A deployment must check peer DMA reachability, memory registration, IOMMU behavior, and the supported driver combination. Otherwise, a nominally compatible software API can still fall back to a different movement path. That fallback may consume host bandwidth and CPU time that were absent from the expected performance budget.

Packet compatibility is also insufficient if the endpoints react differently to congestion. BURST uses observations of the tested commercial RNIC’s connection-management behavior to align on a compatible congestion-control mode. The documented fallback to DCQCN is an implementation finding for the examined environment, not a universal negotiation guarantee across every RNIC firmware.

For retransmission, the main compatible path uses Go-Back-N. More selective retransmission can be enabled when the peer’s behavior is known, but that is a narrower case. Loss, congestion, and receive-not-ready events therefore deserve separate tests. A clean point-to-point throughput result does not establish identical behavior on an oversubscribed or lossy production fabric.

Near-400G throughput is a multi-queue result

The 400G evaluation compares BURST with kernel RXE and kernel TCP on the mixed endpoint path. Native hardware RDMA is also shown as a reference, but that reference uses RNICs at both ends. It is not the same hardware-capability arrangement as the software compatibility problem being solved.

The multi-NIC test uses 1 MB messages, eight queue pairs per NIC, and NUMA-aware placement. With one NIC, the paper reports 387.12 Gbps using three CPU cores. With four NICs, it reports an average of 387.6 Gbps per NIC and 14 CPU cores in total. Those conditions explain how parallel software workers approach the available network rate.

The single-queue-pair result is substantially different. At a 1 MB message size, BURST reaches 121.9 Gbps, while the hardware reference reaches 385.4 Gbps. The appendix attributes the software limit to the processing arrangement that maps a queue to a particular worker and hardware queue. Aggregate NIC throughput cannot be assigned automatically to one application’s single flow.

This distinction is useful for capacity planning. A workload with many independent transfers may expose enough parallelism to use several workers efficiently. A serialized transfer on one queue can remain limited even when other cores and link bandwidth are available. Increasing the line rate does not remove that serialization, and increasing application concurrency can change memory and queueing costs elsewhere.

KV transfer and first-token response are separate outcomes

The application test connects two GPU-equipped nodes in the mixed NIC configuration. Each side has eight GPUs and four 400G NICs, with commercial RNICs on one side. The authors measure KV transfer inside the path to the first output token, reporting an average transfer corresponding to 2.94K tokens of KV state.

Mean KV transfer latency falls from 16.9 ms under kernel TCP to 4.26 ms under BURST. At P99, the corresponding values are 111 ms for TCP and 24.2 ms for BURST. The new mean is approximately 25.2% of the old mean, which is the source of the striking ratio associated with this workload.

The complete first-token metric changes by less: mean TTFT falls from 1,190 ms to 931 ms, a 21.76% reduction. P99 TTFT falls from 3,720 ms to 3,210 ms, about 13.7%. These measurements include work beyond the isolated transfer. The paper’s broad abstract wording should not be used to turn the KV-transfer ratio into a claim about complete inference latency.

The separately reported means also do not provide a complete additive timing decomposition. Subtracting the transfer means does not explain the entire difference in TTFT. Without corresponding per-request traces and a full breakdown, we should not invent the contribution of queueing, overlap, or other stages. The defensible conclusion is that both transfer and first-token response improve in this test, by different amounts.

The evaluated mean KV transfer changes from 16.9 to 4.26 ms, while mean first-token response changes from 1,190 to 931 ms. Separate scales preserve the different measurement scopes. These are independently reported metrics, not additive components reconstructed into one timing stack. Original figure created for this article.

Connection bursts and the deployment decision

Transport bandwidth is only one scaling limit. A restart or expansion can require many connections before useful traffic resumes. BURST implements connection-management handling in userspace and uses host resources to process concurrent setup work while retaining the application-facing interface.

In the reported two-machine, multithreaded connection experiment, the system establishes 24,792 connections per second, approximately 12× the native connection-manager comparison. This is connection creation, not storage operations per second or inference requests per second. It is most relevant when setup bursts are a meaningful part of recovery or scaling time.

We would assess BURST in three separate stages: protocol and memory correctness, transport performance under the application’s queue distribution, and the end-to-end service objective. The last stage should include failures, concurrent workloads, and actual congestion behavior. Saving host CPU in a microbenchmark is useful only if the reserved workers and memory movement fit the host’s broader workload.

The strongest case is a fleet that must keep mixed hardware productive while retaining RDMA-oriented applications. BURST demonstrates that software can bridge that gap with much more performance than the compared kernel paths. Its limits are equally informative: queue parallelism, supported buffer shapes, device access, and non-network inference work still determine how much of the improvement becomes visible to users.

Sources and rights

This independently written analysis uses the final NSDI 2026 paper, including its evaluation and appendices. Measurements retain their original test conditions; deployment recommendations are our interpretation. Original-paper copyright remains with the authors under USENIX publication terms, © 2026. Figures here are newly constructed explanations and data replots, not reproductions of the paper’s visual arrangement.