Buying more memory can be an indirect way of buying bandwidth or isolation. A workload may occupy little capacity but repeatedly stream through its data, while another retains a large working set with relatively few accesses. If the platform allocates memory only as a number of gigabytes tied to vCPUs, those applications cannot ask for the resource combinations they actually need.
RamRyder, an OSDI 2026 collaboration between Samsung Semiconductor and researchers at UC San Diego and Shanghai Jiao Tong University, makes memory channels visible to system software[1]. It controls which channels back a VM’s pages instead of relying entirely on the hardware’s default interleaving. The objective is to separate bandwidth allocation from capacity allocation where the hardware permits, while limiting interference between tenants.
The implementation is a single-host prototype with directly attached CXL memory. It is not a demonstrated multi-host switched pool, and it is not an HBM allocation mechanism for GPUs. For AI infrastructure, its immediate relevance is the CPU-side environment supporting services, graph processing, caching, and memory-intensive jobs. Any extension to other memory domains requires its own topology and isolation mechanisms.
Why idle gigabytes do not describe available service
A memory channel has both storage attached to it and a finite rate at which it can serve requests. The two resources are related physically but need not be consumed in proportion. A VM with cold data may fill memory while barely using its bandwidth. A smaller streaming working set can do the opposite. Consolidating the two looks attractive only if one tenant cannot disrupt the other’s latency.
Default channel interleaving spreads physical addresses across a socket’s channels to exploit parallelism. This is useful for a large application using the socket, but it also means colocated VMs can compete on those channels. Merely assigning disjoint address ranges does not necessarily give them disjoint service resources. A VM with more active cores can generate enough traffic to delay another VM’s requests.
That creates a reason to reserve more hardware than the application’s capacity alone requires. An operator may leave nominally unused bytes unavailable because they are attached to bandwidth reserved for a sensitive workload. Reclaiming those bytes without understanding access rates can destroy the isolation that motivated the reservation. A useful memory-pooling policy must therefore account for both occupancy and access demand.
Moving channel control into the virtualization stack
RamRyder first changes the boot-time memory configuration so that software can identify physical regions belonging to individual channels. It reserves those regions for management through DAX mappings. This replaces the assumption that the entire socket’s memory should always appear as one uniformly interleaved allocation domain.
The runtime then coordinates three layers. A host resource manager assigns memory regions and observes VM demand. A modified QEMU implementation maps the selected regions into the guest. A modified guest kernel uses the exposed topology to place pages across the channels assigned to that VM. The software stack matters as much as the ability to attach CXL devices.
The paper distinguishes channel-level NUMA nodes from the original server-level NUMA organization. The guest can distribute pages across its assigned channels while preserving the broader distinction between local DIMM memory and CXL memory. Applications need not directly manage every channel, but the guest kernel must understand the additional structure.
Running on commodity hardware should therefore not be mistaken for running with an unchanged guest image. BIOS configuration, physical-memory reservation, hypervisor support, and kernel changes are prerequisites. A cloud that cannot modify tenant kernels faces a different deployment problem from an operator controlling a managed VM image.

Baseline isolation and elastic resources are different promises
The initial DIMM allocation remains proportional to the VM’s baseline capacity in the evaluated setup. Those channels provide a more isolated resource assignment. CXL supplies additional capacity and bandwidth that can be allocated dynamically. The paper describes that elastic part as best effort, which is an important distinction from the baseline reservation.
To add capacity without much extra bandwidth demand, the system can place more cold data on a smaller set of CXL channels. To add bandwidth with little additional capacity, it can spread a VM’s memory across more channels. The number of bytes and the number of participating channels become separate control variables, although they remain constrained by actual device capacity and connectivity.
The term channel also needs care. It describes the DRAM-side service resource, including the internal DDR interface of the tested CXL device, rather than simply a PCIe link. The tested CXL devices each expose the relevant single-channel resource. Future or different devices with several internal channels may not provide the same software control. A product’s CXL version alone does not establish this granularity.
Our interpretation is that the system creates a more useful allocation vocabulary, not complete independence between all memory properties. Shared links, controller limits, channel granularity, and page access patterns still constrain achievable bandwidth. An allocation can authorize access to a resource without guaranteeing that the application has enough parallel requests to exploit it.
Page placement determines whether the extra bandwidth is useful
Adding a channel only creates an opportunity for more parallel memory service. Existing pages remain on their previous channels until their placement changes. An allocation-heavy workload may use the new channel through subsequent allocations, but a long-lived working set needs redistribution before its accesses can benefit.
RamRyder uses different cross-tier policies for different demands. Capacity-oriented placement can retain frequently accessed pages in DIMMs and put colder pages in CXL. Bandwidth-oriented placement distributes pages across tiers according to their available service rates. These are different objectives: minimizing access latency for a hot subset is not identical to maximizing aggregate throughput for a streaming working set.
The policy must also reflect the application’s actual locality. Spreading cold pages onto more channels does not accelerate requests that repeatedly hit a small set left elsewhere. Conversely, moving latency-sensitive pages to a slower tier can be counterproductive even if aggregate bandwidth increases. The relevant success metric is application progress or its response-time target, not the number of enabled channels.
Isolation extends beyond the memory channels
Disjoint channels do not eliminate competition in every shared structure. RamRyder also places VM vCPUs on separate core complexes with distinct last-level caches in the evaluated AMD system. This limits cache interference that would otherwise obscure the benefit of separating memory traffic.
The paper’s ablation tests distinguish channel isolation from cache isolation. That is a useful evaluation practice because a faster colocated workload could otherwise be attributed entirely to the channel allocator when CPU placement also changed. Porting the design to another processor requires repeating the analysis with that processor’s cache and core organization.
Small VMs expose the allocation granularity problem. A complete channel may be larger than one tenant’s demand. Sharing it among several small VMs improves utilization but weakens the simple isolation obtained from exclusive channels. The paper discusses combining sharing with throttling rather than claiming that channel partitioning resolves every tenancy size equally well.
An operator must choose whether the sellable unit is a dedicated channel, a best-effort share, or a service-level target enforced by several controls. These are different products. Advertising all three as one elastic-memory guarantee would make it difficult to diagnose which resource was actually promised when interference appears.
Elastic bandwidth has a transition time
The dynamic mechanism adds channel-backed guest memory, updates the guest topology, redistributes pages where needed, and reclaims an equivalent amount from the old allocation. Keeping the final capacity unchanged requires more than toggling a bandwidth setting. The data must reach the channels whose bandwidth is being offered.
In the reported 10 GB read-only test, the observed rates are 38 GB/s before expansion and 68 GB/s after the added CXL channel becomes useful. That transition takes 2.2 seconds; reclaiming the channel takes 1.1 seconds in the same experiment. These are specific transition measurements, not universal provisioning times. Working-set size, access behavior, and migration bandwidth affect the cost.

The resource manager also needs time to observe demand. Its monitoring interval is at least one second in the described approach, followed by page redistribution when necessary. A short burst can finish before the system reacts. The design is therefore better suited to sustained changes than to instantaneous demand spikes that require all bandwidth to be available in advance.
This leads to a practical reserve policy. A latency-sensitive service may need baseline bandwidth sufficient for its fastest expected bursts, using elasticity for longer shifts. Otherwise, the control loop can faithfully detect overload while still reacting too late to protect the requests that triggered it. Prediction could help, but it is a future improvement rather than an established result here.
Separating hardware experiments from trace consolidation
The hardware evaluation uses an AMD EPYC Zen 5 server and Samsung CXL 2.0 memory devices. Its four-VM setup partitions populated DIMM channels and uses workloads including memory microbenchmarks, Redis, Memcached, STREAM, and graph processing. These experiments examine interference and application behavior under controlled resource allocations.
The ideal baseline gives the workload exclusive hardware access. Shared and hardware-throttled configurations provide other comparison points. Results close to the ideal baseline show that the chosen isolation and placement work for those tests; they do not establish a universal guarantee across arbitrary applications, processors, or CXL devices.
The cluster-level utilization analysis is a different kind of evidence. It pairs workloads from cloud traces when their combined capacity and bandwidth demands fit the resource limits at each recorded time. Selected pairs are also replayed using a custom generator. This demonstrates an opportunity for complementary placement, not an observed production rollout across a cloud fleet.
The headline utilization improvements need their percentile context: the paper’s detailed discussion associates average-capacity improvement with P30 and average-bandwidth improvement with P90. We do not combine them into a single claim about a universal cluster-average gain. A real scheduler must discover compatible workloads without future knowledge, preserve other resource constraints, and pay any movement cost.
The deployment decision for a managed AI platform
Start by measuring per-VM bandwidth demand and latency sensitivity separately from capacity occupancy. A capacity-heavy cache and a bandwidth-heavy preprocessing task may complement each other, but CPU, cache, and I/O interference can still make the pair unsuitable. Consolidation should be evaluated at the application target, including tail latency, rather than only at the host’s average utilization.
Next, verify the hardware control surface. Can the firmware expose channel-specific regions? Can the guest and hypervisor represent them without breaking existing memory-management assumptions? Does the CXL device expose the needed allocation granularity? These questions determine whether the prototype’s mechanism can be reproduced on the intended platform before performance comparisons become meaningful.
Finally, test the transition rather than only the final allocation. Grow and shrink bandwidth while a realistic working set is active, measure page-fault and migration effects, and include bursts shorter than the monitoring interval. Record how much baseline capacity must remain reserved to satisfy the service target throughout the change.
RamRyder shows that memory bandwidth can become an explicit software-managed resource when topology and page placement cooperate. Its most useful lesson is not that memory becomes fully fluid. It is that the operator can separate some previously bundled demands, provided it also accepts the granularity, guest-software requirements, and transition time of the mechanism that makes that separation possible.
Source and copyright note
This independent editorial analysis uses the final OSDI 2026 paper[1]. Reported hardware measurements and trace results retain their separate scopes; deployment recommendations are our interpretation. The authors represent Samsung Semiconductor alongside UC San Diego and Shanghai Jiao Tong University. Text and explanatory figures were created for Silicon & Systems without reproducing source figures or tables. Original-paper copyright remains with its respective rights holders, © 2026; proceedings publisher: USENIX Association.