A cloud can have enough network bandwidth and still run out of networking capacity. A newly created connection needs a policy decision and a state entry. An established connection needs that entry to remain accessible at packet rate. An accelerator introduced to perform those operations needs power, cooling, cabling, and a place to install it. None of these requirements disappears when the forwarding chip gains more ports.
Microsoft’s SONiC DASH SmartSwitch addresses these constraints together[1]. It places data processing units (DPUs) inside a switch already needed by the datacenter and narrows the service programming model enough to make hardware implementation predictable. The important change is not simply a more capable switch. It is a different unit for purchasing, installing, programming, and maintaining stateful network services.
This is infrastructure for virtual networks, gateways, and load balancing, not a replacement for a GPU collective library or a new scale-up interconnect. Its relevance to AI lies in the service environment around accelerators: tenant isolation, private endpoints, inference backends, and the network functions connecting compute to storage. Applying its results to training all-reduce bandwidth would confuse two distinct workloads.
Why the earlier DPU pool remained difficult to install
Moving network processing away from customer CPUs can release compute capacity, but the destination matters. A dedicated DPU appliance creates a shared pool without requiring every host to carry the same acceleration resources. It also introduces equipment whose only purpose is to host and connect those DPUs. The pool may be logically elastic while being physically awkward to expand.
In the preceding Azure design, Sirius, additional servers and switches supported the DPU pool. Existing compute rows did not necessarily contain spare rack positions for that appliance. Placing another rack beside a row affected circulation and maintenance access; removing a compute rack sacrificed sellable capacity. A design that looks attractive when measured only in packets per second can therefore fail a facility-level deployment test.
SmartSwitch instead occupies the T1 switching tier (the middle-of-rack role in the paper’s Azure topology). The chassis combines ordinary packet forwarding with modules for stateful processing. Power distribution, cooling, and network attachment are shared with switching equipment that the datacenter already needs. This integration reduces auxiliary equipment; it does not make the DPU resources free or eliminate their service lifecycle.
A restricted pipeline with useful configuration
The software decision is equally consequential. A general virtual switch can expose arbitrary rule shapes, actions, and chains of processing layers. That flexibility is convenient for a software implementation, but hardware must then support combinations whose storage, lookup, and execution requirements are difficult to bound. Offloading only the common established-flow path leaves expensive new-flow handling behind.
DASH uses a predefined sequence of functional stages instead. The stage count, connectivity, match-key structure, supported actions, and configurable attributes have specified meanings. An operator can populate tables or select supported fields without redefining the hardware’s matching machinery. Disabling a field in an allowed key is different from asking the device to implement a new match type.
The paper’s pipeline contains 13 stages, derived from recurring cloud-service requirements. This is not a claim that every conceivable network function fits those stages. It is a choice to support a known service family well. Adding an incompatible function remains an architectural change requiring a revised specification and vendor implementation, rather than an ordinary tenant configuration update.
Consider the distinction between a virtual-network mapping and address translation at a gateway. Their policies differ, but they still require identifying the tenant context, finding a connection, determining transformations, and applying those transformations consistently. A shared set of primitives can express these recurring operations without exposing the unrestricted behavior of a general software switch.
Keeping connection state beside the work that modifies it
SmartSwitch’s NPU is a network processing unit, not a neural-network accelerator. It handles the high-bandwidth underlay forwarding role. DPUs perform the stateful DASH service processing, including both the established-flow path and the work needed when a flow is first encountered. Treating the NPU and DPU as interchangeable accelerators would miss the memory and update-rate rationale.
When a packet belongs to an existing connection, the DPU can look up the associated state and apply its recorded actions. A miss requires deriving policy results and creating state before subsequent packets can use the accelerated path. On the evaluated implementation, DPU CPU software participates in state insertion, while the DPU’s packet-processing hardware handles established flows. Keeping these functions together avoids turning every state update into a coordination problem across separate device classes.
Tenant traffic reaches the appropriate service instance through encapsulation and endpoint-aware forwarding. The physical location of an endpoint’s state need not be exposed to the VM. However, that indirection still requires a placement policy and sufficient connectivity to the selected DPU. A pool’s aggregate free capacity cannot cure a single overloaded endpoint unless the placement and migration mechanisms can actually use it.

Three capacities that should never become one headline
The tested Cisco platform combines a 12.8 Tbps forwarding chip with eight 200 Gbps DPUs, arranged in four replaceable sleds. Its external network-facing port capacity is 11.2 Tbps; internal DPU connectivity accounts for a separate 1.6 Tbps. These quantities describe different paths. Advertising the full forwarding-chip capacity as stateful service throughput would overstate what the DPU pool supplies.
The large-packet experiment reports 1.53 Tbps with all eight DPUs. Small packets reach a different limit: approximately 391 million packets per second. A bandwidth result cannot substitute for a packet-rate test, because shortening packets increases the number of lookups and per-packet operations required to deliver the same bit rate. This distinction matters for short RPCs and control traffic even when bulk transfers dominate byte volume.
Connection creation introduces another limit. The eight-DPU configuration reaches 19.22 million new connections per second without background flow entries, compared with 16.82 million when each DPU already holds 32 million background entries. The latter test represents 256 million retained entries across the chassis. It demonstrates why a connection-rate claim should be accompanied by the occupancy of the state table rather than presented as an unconditional constant.
The load is distributed evenly across DPUs in the scaling experiments. That condition supports the hardware scaling result but is not a guarantee for skewed customer traffic. A deployment must measure the busiest DPU and the most demanding endpoint, not only divide total traffic by the number of modules. Average packet latency measured with one DPU is also not an application-level tail-latency guarantee.

Total power and incremental power answer different questions
A facility operator deciding whether an existing row can accept the system needs incremental power and space. A buyer comparing complete implementations may instead need total consumption per delivered unit of service. Those are both valid measurements, but they cannot be exchanged midway through an efficiency argument.
The full-load SmartSwitch measurement is 1,040.83 W. Relative to the ordinary T1 equipment used as the infrastructure baseline, its additional consumption is 448.72 W. Similarly, the chassis occupies 2U, while the additional rack-space requirement is 1U. The smaller incremental figures reflect equipment being replaced, not a lower total power draw or a physically smaller chassis.
The comparison with Sirius normalizes to half of its redundant appliance. Reported improvements of 1.81× in power efficiency and 2.68× in rack-space efficiency use packet throughput per unit of the respective resource. They do not establish an equivalent reduction in whole-datacenter electricity or an equivalent increase in usable floor area. Procurement must also restore the redundancy, reserved capacity, and maintenance requirements appropriate to the service being purchased.
Our interpretation is that integration creates value by removing support equipment and fitting an existing deployment pattern. The amount realized by another operator depends on its starting topology. A newly built facility with abundant service-appliance space and an existing row with no available rack position face different constraints even if their traffic is identical.
P4 as a behavioral reference, not a universal binary
An open API alone does not establish that two devices implement the same service. Subtle differences in packet transformation order or state updates can change tenant-visible behavior despite similar function names. DASH addresses this with a P4 reference implementation of the expected pipeline semantics, backed by tests and generated control interfaces.
Vendors remain free to implement equivalent behavior using their own hardware mechanisms. The reference model is therefore not a promise that a single P4 program compiles unchanged onto every DPU. It makes the required behavior inspectable and executable before a specific hardware implementation arrives. Functional testing can begin against the software model and later be applied through the same control interface to devices.
This moves part of the switching cost between suppliers from service reimplementation toward conformance verification. That is valuable, but functional agreement does not imply equal table capacity, connection rate, fault behavior, or operational tooling. Those properties still need target-specific tests. A vendor-neutral interface is strongest when paired with explicit limits, not when it hides them.
Maintenance changes the amount of sellable capacity
Combining services in one chassis creates an obvious operational question: what happens when it must be updated? Rebooting the whole device removes both forwarding and the attached service pool. The paper describes updating DPUs in sequence so maintenance reduces available service capacity rather than taking the entire pool down at once.
That distinction only helps if enough headroom remains. An operator should evaluate peak demand against the capacity available while a module is unavailable, including the cost of moving or synchronizing state. Selling the entire benchmark capacity and then expecting a seamless rolling update is not a defensible availability plan. The relevant reserve depends on failure and maintenance procedures, not just nominal DPU count.
State synchronization also consumes real network resources. The Azure deployment favors T0 paths for bulk synchronization between T1 SmartSwitches, balancing congestion concerns against correlated failure exposure. This is a topology-specific operating decision. Another network should repeat that analysis with its own oversubscription and failure domains rather than copying the tier choice as a universal rule.
Debugging becomes harder as state changes faster. Dumping a large flow table can yield a view that is already stale when collection ends. The reported response is selective flow inspection and tracing which stage entries a packet actually hits. These capabilities matter because a faulty policy or missing entry can look like an ordinary network timeout to an application owner.
The adoption test for an AI service network
An AI platform should first identify which traffic needs these stateful services. Large training transfers, storage access through private endpoints, public inference ingress, and internal agent RPCs need not follow the same path or require the same state. The benefit of a shared service pool depends on the traffic that actually traverses it, not the total number of GPUs installed.
Next, separate steady-state traffic from churn. A stable set of high-bandwidth flows stresses a different resource from short-lived connections created by a burst of inference workers. Measure bit rate, packet rate, state occupancy, and connection creation together under realistic tenant skew. Repeat the test with the planned maintenance reserve removed and with state synchronization active.
Finally, ask whether the service semantics are sufficiently stable to justify a constrained hardware model. DASH is attractive when recurring operations dominate and multiple suppliers can implement a common behavioral specification. A rapidly changing, application-specific packet-processing function may still need software flexibility. The design decision is not hardware versus software in the abstract, but which behaviors are mature enough to standardize and accelerate.
SONiC DASH’s lasting contribution is the alignment of that software decision with a physically installable system. Its performance results establish a capable stateful service path; its deployment lessons explain why that path can become usable cloud capacity. The purchasing unit should therefore include maintenance headroom, observable state, and verified service behavior alongside ports and DPUs.
Source and copyright note
This independent editorial review is based on the final NSDI 2026 paper[1]. The source authors are affiliated with Microsoft, Microsoft Research Asia, and The Chinese University of Hong Kong. Reported measurements are attributed to their testbed and deployment; adoption criteria and system-level interpretations above are our analysis. Text and explanatory figures were created for Silicon & Systems without reproducing the paper’s tables or figures. Original-paper copyright remains with the respective rights holders, © 2026; the proceedings are published by USENIX Association.