NIC receive ring은 queue 하나처럼 보이지만 두 producer-consumer relation을 구현합니다. Software는 NIC가 consume할 empty buffer를 descriptor에 둡니다. Packet이 도착하면 NIC가 software가 consume할 full buffer를 돌려 줍니다. 두 방향을 하나의 circular order에 묶으면 NIC와 CPU가 ring 전체를 진행하기 전에는 먼저 반환된 empty buffer도 재사용할 수 없습니다.
많은 100Gb/s adapter의 default ring은 burst를 흡수하도록 1,024 entry를 가집니다. 1,500-byte buffer라면 core마다 약 1.5MiB I/O working set을 reserve합니다. Core가 많으면 last-level cache와 DDIO available portion을 넘습니다. New DMA write가 unprocessed packet을 evict하고 CPU가 memory에서 다시 읽으며 bandwidth가 늘고 slow core는 더 뒤처집니다[1].
Small ring은 cache에 맞지만 큰 burst를 drop합니다. Core 사이에 ring 하나를 공유하면 buffer는 줄지만 delivery progress가 결합됩니다. Overloaded core가 common ring을 채워 idle core로 갈 packet도 막습니다. RxBisect는 buffer capacity는 share하되 packet delivery는 독립적으로 유지해야 한다는 주장입니다.
Allocation ring과 reception ring의 분리
각 core는 empty buffer를 담은 allocation(Ax) ring을 제공합니다. Bisected reception(Bx) ring은 processing core로 전달된 packet notification을 담습니다. Bx ring 하나가 여러 Ax ring에서 buffer를 가져올 수 있습니다. Burst 중 preferred allocator가 비면 NIC는 연결된 다른 Ax ring의 buffer를 사용하면서 originally selected Bx consumer에 그대로 알립니다.
Receive-side scaling은 processing core를 계속 고릅니다. Buffer ownership만 이동합니다. Empty buffer union은 one shared packet queue의 lock 없이 shared resource가 됩니다. Software는 small per-core Ax ring을 쓰거나 allocation core 몇 개를 dedicate할 수 있습니다. Bx descriptor는 permanently reserved packet-buffer set을 가리키지 않으므로 notification burst를 위해 더 크게 둘 수 있습니다.

선택한 buffer가 다른 allocation ring에 있으면 hardware가 descriptor operation 하나를 더 수행하지만 packet 자료는 같은 DDIO-local memory에 기록됩니다. 원문은 dependent DMA critical path가 existing private 및 shared design과 같다고 봅니다. Evaluated imbalance extreme에서 allocator traffic은 cycle의 0.2% 미만이었습니다.
실제 sizing constraint인 cache capacity
22MiB LLC를 가진 16-core processor와 100Gb/s ConnectX-5 NIC 두 개에서 private-ring size를 늘리면 measured throughput은 최대 20%, latency는 최대 37배 나빠지고 memory bandwidth는 최대 4.9배 증가했습니다. Ring size 128 이하로 working set이 DDIO way에 맞을 때 line rate를 유지했습니다. 그 뒤 buffer가 DDIO capacity, full LLC를 차례로 넘을 때 result가 두 단계로 떨어졌습니다.
Core를 줄이는 것은 일반 해법이 아닙니다. Large packet은 core 16개를 모두 쓰기 전에 peak throughput에 도달했지만 64-byte packet은 모든 core가 필요했습니다. Changing traffic에 맞춰 processing parallelism과 burst absorption을 제공하면서 maximum core count와 maximum ring depth를 곱한 buffer를 영구 reserve하지 않아야 합니다.
RxBisect와 shared-ring design은 private 1,024-entry ring의 no-drop throughput을 buffer working set 8분의 1로 맞췄습니다. Shared ring과 달리 rxBisect에서는 slow consumer가 buffer pool을 점유하지 않습니다. Synthetic skew에서 target core의 processing이 느려지면 shared-ring throughput은 최대 60% 떨어졌지만 rxBisect는 software emulator가 limit이 될 때까지 line rate를 유지했습니다.
Imbalance claim을 시험한 real trace
평가는 NAT, load-balancer network function, MICA key-value store를 사용했습니다. Server에는 16-core Xeon Silver 4216 두 개, 22MiB LLC, back-to-back 100Gb/s ConnectX-5 NIC pair 두 개가 있었습니다. Commodity NIC가 proposed interface를 제공하지 않아 RxBisect는 software로 emulate했습니다.
Imbalanced CAIDA trace와 co-located PageRank에서 rxBisect는 idealized dynamic shared-ring policy보다 load balancing 16%, NAT 20% 높은 throughput을 보였습니다. STREAM memory pressure를 더해도 emulated advantage는 최대 16%였습니다. Low traffic에서는 shared-ring synchronization이 두 function에서 private ring보다 packet당 cycle을 최대 34%, 46% 더 사용했지만 rxBisect는 shared-tail lock을 피했습니다.
Ordinary private ring 대비 reported throughput은 최대 37%, latency는 최대 11배 개선됐습니다. RxBisect가 line rate를 지키고 기준이 packet을 queue한 overload transition에서 큰 latency ratio가 나왔습니다. Constant per-packet acceleration으로 읽으면 안 됩니다.
Emulation이 만드는 조건
저자는 emulator를 다른 NUMA node에 두고 packet buffer와 worker는 NIC-local로 유지하며 doorbell과 extra Bx DMA를 재현했습니다. Existing ring scheme에서 emulation은 throughput을 최대 12% 낮추고 latency를 최대 94% 높여 conservative comparison이라는 근거를 줍니다. 그래도 real NIC에는 software가 재현하지 못한 pipeline, cache-coherence, descriptor-fetch, firmware constraint가 있을 수 있습니다.
Native adoption은 driver, DPDK, firmware, silicon 사이 ABI를 바꿉니다. NIC는 nonempty Ax ring을 고르고 어느 buffer를 consume했는지 보고하며 isolation을 보존하고 reset 뒤 복구해야 합니다. Tenant가 protection domain 밖 memory를 빌리지 못하도록 queue configuration limit도 필요합니다.
Allocation과 delivery가 독립적으로 진행되면 descriptor 추적도 정확성의 일부가 됩니다. Software는 반환된 buffer를 local allocator에 돌려야 하는지 다른 core의 cache를 거쳐 release해야 하는지 알아야 합니다. Ax와 Bx head가 서로 다른 속도로 wrap해도 NIC가 같은 entry를 두 번 consume해서는 안 됩니다. 연결된 Ax ring이 모두 비었을 때 packet을 drop, pause, redirect하는 조건도 정의해야 하며, 계수기는 receive congestion과 allocation starvation을 구분해야 합니다.
NUMA placement는 cache 이점을 뒤집을 수 있습니다. 같은 socket 안에서 buffer를 공유하면 의도한 LLC locality를 유지하지만 다른 socket의 Ax ring에서 빌리면 cache pressure 대신 interconnect traffic과 remote-memory latency가 생길 수 있습니다. Association set은 NUMA domain 안으로 제한하고 local capacity가 소진될 때만 확장하는 hierarchy가 필요합니다. 이렇게 해야 일반 경로의 locality와 심하게 치우친 burst를 흡수할 emergency pool을 함께 유지할 수 있습니다.
Interface는 load balancing을 대체하지 않습니다. RSS 또는 application-aware steering이 어느 core가 flow를 처리할지 정합니다. RxBisect는 그 core의 temporary buffer shortage가 다른 곳의 spare capacity를 막지 않게 합니다. LLC partitioning 및 header-only DDIO scheme과도 complementary합니다. 이들은 interference를 줄이고 split interface는 buffer set 자체를 줄입니다.
Coupled resource를 분리하는 판단
Conventional ring은 burst capacity, memory ownership, processing order를 묶었습니다. Lower line rate에서 queue 하나가 core 하나를 처리할 때는 합리적이었습니다. 수백 Gb/s에서는 core별 provisioning이 descriptor depth를 cache-capacity cost로 바꿉니다. Whole queue sharing은 cost를 없애지만 slow consumer의 backpressure를 퍼뜨립니다.
RxBisect는 fungible resource인 empty packet buffer만 공유합니다. Scheduling resource인 delivery queue는 독립적으로 유지합니다. Storage submission queue와 accelerator command ring에도 같은 질문을 적용할 수 있습니다. Circular structure 하나가 한 방향의 capacity와 반대 방향의 completed work를 함께 나르면 역할 분리가 locality와 imbalance tolerance를 개선할 수 있습니다.
Deployment decision은 core별 working-set byte, DDIO way, burst-loss tolerance, flow skew, allocator crossing, memory bandwidth, saturation 근처 tail latency를 측정해야 합니다. Proposed number는 hardware exploration을 정당화하지만 immediate replacement를 증명하지 않습니다. Native prototype이 extra selection 및 notification logic으로 line rate, isolation, reset semantics, driver simplicity를 보존하는지 보여야 합니다.
Mixed packet size와 changing flow affinity를 long interval에서도 시험해야 합니다. 평균적으로 균형 잡힌 traffic은 모든 receive interface에 가장 쉬운 조건이며 rxBisect의 가치는 core별 processing time과 arrival rate가 달라질 때 나타납니다. Ax depletion, cross-core buffer borrowing, Bx occupancy, DDIO miss, allocator handoff를 보여 주는 계수기가 있어야 claimed resilience를 운영 중 진단할 수 있습니다. 이를 통해 throughput 하락이 compute saturation, buffer capacity 고갈, 불리한 NUMA association 중 어디에서 왔는지 구분할 수 있습니다.
출처와 저작권 안내
이 글은 Silicon & Systems가 작성한 편집 분석으로, interface와 measurement, limitation을 우리 표현으로 다시 썼습니다. 원문의 문장, 표, 도판은 재수록하지 않았고 도판은 이 글을 위해 새로 만들었습니다. 전체 논문은 USENIX OSDI 2025 발표 페이지에서 확인할 수 있습니다. 저작권은 저자에게 있습니다. 2025.