A large storage pool can gain capacity without gaining a proportional ability to retrieve that capacity. For hard disks, moving to the requested location consumes time before useful data transfer begins. Splitting one moderate-sized read across many disks can multiply that positioning work even when the file’s protection scheme is space-efficient.

Okapi separates two decisions that many cluster file systems tie together: where successive file data is striped and which blocks are protected by the same erasure code[1]. The Carnegie Mellon University and Google collaboration implements the idea in HDFS and evaluates the resulting prototype on a controlled HDD cluster.

This is relevant to AI infrastructure’s capacity-oriented storage tiers, where training corpora, retained checkpoints, and shared datasets may outlive their fastest cache copies. However, the paper is not an AI-training benchmark. Its findings should inform storage design without being converted into an unmeasured GPU-utilization gain.

Two widths that answer different questions

Striping distributes successive portions of a file across data blocks on different devices. Its width controls how many devices participate in an access and how much work each receives. Wider striping can expose parallel bandwidth, but it can also turn one request into several smaller disk operations.

Erasure-code grouping answers a different question. A group contains data blocks and additional parity blocks, allowing missing data to be reconstructed after failures. Increasing the data width for a given parity count can improve capacity efficiency, while changing reconstruction traffic and the reliability calculation.

When a system forces the stripe width to equal the data width of the code, selecting one configuration simultaneously selects both behaviors. An application that wants a relatively concentrated read layout but a space-efficient wide code cannot express that combination. Choosing a narrower code merely to change reads can increase parity overhead unnecessarily.

Okapi makes these settings independent per file. Data remains striped in an ordered sequence, while consecutive sets of data blocks form protection groups that need not align with stripe boundaries. It does not invent a new error-correcting code; it changes how placement and protection are composed.

Why positioning work changes the best layout

For an HDD, reading a small amount from several disks incurs several positioning operations. Reading larger contiguous portions from fewer disks can amortize that fixed work over more useful bytes. Under load, reducing aggregate disk work can improve cluster throughput even if an isolated request uses less parallelism.

At low load, wider parallelism can still shorten a particular request. Large sequential reads can also benefit from more participating disks. The best stripe width therefore depends on request size, concurrency, and the performance objective, not simply a preference for narrow layouts.

The paper reports Google observations that motivate per-file tuning: many sampled files retain a stable read size even while their protection scheme changes. In the studied clusters, 64–94% of sampled files used the same read size over 150 days. That is evidence about those samples, not a claim that every application or future AI dataset has a stable access pattern.

A conceptual HDD storage view distinguishes device positioning work from useful transfer, while the deterministic overlay separates read width from erasure-code grouping. The generic material is not the evaluated cluster’s product photograph. Original figure created for this article.

Avoiding a second complete metadata map

Separating stripes and groups could require two large location maps. Okapi avoids that duplication by making group membership derivable from the ordered data blocks and the configured widths. The system can determine which group contains a block without storing an independent arbitrary grouping structure.

The HDFS implementation separates structures for data stripes and parity groups while reusing established block-management mechanisms. Normal reads use the data-stripe mapping and need not consult parity groups. Reconstruction and protection changes derive the additional relationships when necessary.

This constrains the flexibility in a useful way. Okapi does not support completely arbitrary combinations of blocks simply to maximize theoretical placement freedom. Its regular grouping rule preserves compact metadata and predictable lookup behavior, making the additional control more practical inside an existing file system.

The approach also keeps the durability requirement on physical placement. Blocks belonging to a protection group must occupy distinct failure domains. Logical independence between widths cannot override the fact that a shared failed device may remove several blocks at once.

Partial parity instead of retaining every unfinished group

During sequential file creation, a protection group can span data from several stripes. Waiting until all of that data is available could force the client to buffer a large amount of unfinished content. Decoupling would then shift storage efficiency into excessive client-memory consumption.

Okapi computes partial parity contributions as data arrives and retains those contributions instead of all original data. Because the code’s operations are linear, contributions can be combined into the final parity. Only completed parity is written out as the final protection data.

The distinction between buffering and durability is important. The prototype retains the comparison system’s write-completion semantics, including recovery of incompletely protected groups after a client failure. It does not establish immediate synchronous parity durability for every arriving cell.

Applications requiring stronger synchronous-write guarantees would need additional mechanisms and costs. This matters when applying the idea to checkpoints: an application should not delete its only previous recovery point merely because a data transfer has started. The exact event that makes the new checkpoint durable belongs in the end-to-end protocol.

Failure reads can require different data

Normal reads follow the stripe layout. If a requested block is unavailable, reconstruction follows the protection group instead. When the two widths differ, some data needed for recovery may lie outside the portion the client was otherwise going to fetch.

Okapi caches useful data during degraded reads to avoid fetching it repeatedly, especially for large sequential accesses. However, this does not remove every unfavorable case. Mid-sized requests can incur more reconstruction traffic than the coupled layout, and tail latency may worsen under failures.

In one measured 24 MB case with 6-of-9 protection and three-wide striping, degraded-read latency is 33% worse. Such degraded reads occur less often because normal requests touch fewer disks, but that frequency argument does not make an individual affected request faster.

An operator therefore needs both the distribution of failure exposure and the latency of requests that encounter it. A throughput improvement in normal operation should not silently become a guarantee for recovery mode, particularly when checkpoint restoration or another urgent workflow must proceed during a device failure.

Changing protection without rewriting every data block

A coupled layout may need to rewrite file data when the erasure-code width changes, because the new code width also dictates a new stripe layout. This can create large background traffic precisely when a failure-rate increase makes protection changes urgent.

Okapi keeps the data stripe layout and constructs new protection groups around the existing ordered blocks. It reads the data needed for new parity and writes the new parity, avoiding the ordinary need to rewrite all data merely to change grouping.

Avoiding a complete rewrite does not mean zero data movement. New groups can contain blocks that share a failure domain. Those collisions require relocation unless the original placement anticipated the relevant future groupings. The paper’s longer-term simulation includes such relocation costs.

The transition also needs a safe completion point. The implementation retains old parity until new protection is complete, then changes the metadata atomically. Saving I/O is useful only if a concurrent failure does not expose the file between two incomplete protection schemes.

What was measured on actual disks

The prototype evaluation uses twenty machines: one metadata node and nineteen data nodes. Each data node has a 1 TB, 7,200-rpm HDD, with a 40GbE cluster network. The file-system tests use 8 MB data blocks and 1 MB striping cells.

These details define the bottleneck being studied. Mechanical positioning costs are central to the observed benefit. NVMe devices remove that particular mechanical cost, although request fan-out and protection changes can still matter. The paper does not measure an equivalent percentage gain for an all-flash training store.

Microbenchmarks using the same 6-of-9 protection show read-throughput improvements up to 80% and seek-rate reductions up to 70% for selected access sizes and tailored stripes. The same code preserves the nominal protection and capacity overhead, so the comparison is not simply trading away parity to obtain more speed.

The maximum remains workload-dependent. When the tailored stripe width equals the coupled layout’s width, the design should not be expected to create a new source of bandwidth. Its benefit comes from choosing a previously unavailable combination, not from making an unchanged disk intrinsically faster.

A realistic distribution, not a production rollout

A separate experiment draws read sizes from a Google production distribution and executes a synthetic read-only workload on the testbed. Sixty-four client threads each perform 10,000 reads, with the same protection configuration used for the compared systems.

That experiment reports 55% higher sustained throughput, 65% fewer total seeks, and a 35% lower seek rate. Total seeks and seeks per second have different denominators because the workload finishes sooner. The reported 36% reduction concerns end-to-end run completion time, not a universal 36% decrease in each request’s latency.

These are useful results for a distribution dominated by reads that benefit from concentrated access. They do not demonstrate that Google deployed Okapi or that a mixed workload with frequent updates will see the same behavior. The distinction between production observations and prototype evaluation must remain explicit.

The Google-derived synthetic read-only test reports higher throughput and a shorter complete run, while total seek count and seek rate decrease by different amounts. Each result is normalized to the coupled HDFS baseline for the same workload. Original figure created for this article.

Metadata overhead depends on what is being counted

The paper’s small overall memory overhead should not be read as negligible growth in every metadata structure. In the reported heap accounting, the combined allocation share rises from 3% to 3.74%, a difference of 0.74 percentage points. The per-file structures themselves show larger relative increases.

For the examined configuration, the reported per-file growth is 26% in the block mapping and 22% in the inode mapping. Both can coexist with a small change in total heap share because those structures occupy only part of the process’s allocation.

The operational implication is to measure the relevant limit. A metadata server constrained by these maps may care more about their per-file growth than a percentage of a generously provisioned heap. File size, group width, and object counts should accompany any memory-capacity estimate.

Transition savings are not all the same experiment

Basic regrouping avoids data rewrites and approaches a 50% reduction in the examined transition I/O comparisons. Combining decoupling with techniques that also reduce parity-recomputation reads can produce larger savings, up to the reported 70%. That combined maximum should not be attributed to every ordinary regrouping operation.

The emergency-transition example is an analytical scenario informed by a Google incident, using representative rather than fully disclosed fleet numbers. The six-year disk-adaptive evaluation uses Backblaze failure data in a simulation and reports approximately 38% less transition I/O, including placement repairs.

These analyses broaden the argument beyond a short disk benchmark, but they remain different forms of evidence. They support the proposition that protection changes consume meaningful resources and that avoiding unnecessary rewriting helps. They do not supply a measured universal recovery deadline for a live exascale service.

Implications for AI data infrastructure

For a capacity tier behind AI workloads, a useful first step is to measure per-file access sizes and whether they remain stable across jobs. A training pipeline that reads fixed-size shards, an interactive retrieval service, and checkpoint restoration may prefer different layouts even when their durability target is identical.

Tests should include normal and degraded reads under representative load, not only peak sequential bandwidth. The selected width may optimize throughput while worsening the far tail during failures. Choosing between those outcomes is an application decision that the storage system should expose rather than hide inside the erasure-code setting.

Protection transitions should also be evaluated alongside foreground reads and emergency recovery. The capacity saved by a wider code is less attractive if changing to it consumes the I/O reserve needed to keep training supplied. Conversely, retaining a read layout while adjusting protection can make the file’s performance more predictable over its lifetime.

We believe Okapi’s durable insight is the separation of objectives. Read placement should answer how the application accesses data, while protection grouping should answer how the service tolerates failures and pays for redundancy. The two still interact through physical devices and recovery, but forcing them to share one width discards useful choices before the operator can evaluate them.

Sources and rights

This independently written analysis reviews the OSDI 2025 paper. Experimental and observational findings are attributed to its authors; AI-infrastructure implications are our interpretation. Original paper copyright remains with the authors, © 2025. Illustrations were created for this article, with conceptual hardware material clearly distinguished from photographs of the prototype.