A training server can wait for data even when the storage devices have unused bandwidth. Opening a sample requires finding the file and checking the directories on its path before the application can read its contents. When metadata is distributed across servers, those steps can produce several network requests for one useful file operation. Adding SSD bandwidth does not remove that work.

FalconFS, developed by Huawei and Shanghai Jiao Tong University, changes where path resolution happens[1]. It targets pipelines with many small files, shuffled training access, and bursts of activity within directories. Instead of making every training client retain enough directory metadata to avoid remote lookups, it gives metadata servers the directory information needed to resolve paths themselves.

The adoption question is therefore not simply whether FalconFS wins a storage benchmark. It is whether directory lookup amplification consumes a meaningful share of the input pipeline, whether server-side replication fits the namespace, and whether the application’s filesystem semantics fit the implementation. The final paper provides substantial design detail, but its benchmark and production claims need to remain separate.

The cache that competes with the data loader

Client-side metadata caching works well when recently used paths are likely to be reused. Training can create a less favorable access pattern: each epoch traverses a large dataset in a new random order. Near-root directories remain useful, but lower directory entries may be evicted long before their next access. A large file count is not automatically a cache problem; the important quantities are the directory working set, access order, and memory budget.

The client memory used for that cache is not spare capacity. Data augmentation, decoded samples, prefetch queues, and intermediate CPU results also need memory. Increasing the directory cache can reduce metadata traffic while reducing the space available to overlap input preparation with accelerator execution. Looking only at filesystem throughput hides that competition within the training host.

FalconFS’s motivating observation is that many clients share relatively few metadata servers. Storing a compact directory representation at those servers can amortize the information across clients rather than duplicate a larger kernel representation everywhere. This shifts the memory burden; it does not eliminate the namespace or make its size irrelevant. A deployment must account for each metadata server’s replica and growth over time.

Replicate the directories, partition the files

The design separates directory information needed for path resolution from individual file metadata. Each metadata node can maintain the directory tree, while file inodes are partitioned across nodes. File contents reside in the storage layer. Replicating the namespace therefore does not mean copying every file’s data or every file inode onto every metadata server.

This distinction is central to scalability. A request reaching the node that owns a file can check its path using local directory information and then operate on the local inode. The client no longer needs to walk the entire remote directory chain first. For the common case, this removes a sequence of client-server lookups before the actual operation.

The namespace replicas are maintained lazily and need not already contain every newly created directory. A node encountering missing information retrieves it from the relevant owner. Consequently, one-hop access is a common-case design goal, not a promise for every operation. Cold paths, exceptional placement, and changing metadata can still require additional communication.

FalconFS separates two metadata populations: directory information is available at metadata nodes for local path checks, while file inodes are partitioned. Clients send the operation with its full path rather than performing the remote directory walk themselves. The illustration is a newly constructed explanation, not the paper’s layout. Original figure created for this article.

Why hashing only the filename needs exceptions

Sending a request to the correct server is harder than removing a cache. Hashing a complete path makes placement simple, but renaming a directory can then change the placement of its entire subtree. Deriving placement from parent identifiers avoids that particular problem but may require resolving the parent before the client knows where to send the request.

FalconFS normally places file metadata using the filename. Large directories with varied names can distribute their files across metadata servers, reducing the tendency for one directory burst to overload one server. A directory rename does not rewrite every descendant’s filename, which avoids the subtree-wide relocation implied by full-path hashing.

However, filename distributions are not uniformly benign. Many directories may contain a file with the same conventional name. FalconFS maintains exceptions that either override placement for a name or incorporate a parent-directory identifier. The latter can require forwarding through another metadata node. Clients also keep exception information, so the term stateless should not be read as literally zero local state.

This reveals an important distinction between balanced storage and balanced activity. The coordinator’s inode-distribution policy can spread stored metadata, but a newly hot file or a synchronized access burst still needs operational attention. The deployment test should examine request rates and latency by node, not only the number of inodes assigned to each node.

Lazy replication without skipping permission checks

Directory creation and directory removal have different synchronization requirements. A newly created directory can be fetched by another node when needed. Removing or renaming a directory, or changing permissions, must prevent another node from using stale information to authorize an operation. FalconFS uses coordinated invalidation and locking for those changes.

The practical benefit is avoiding eager synchronization for every common creation. The cost is that destructive or permission-changing operations involve broader coordination. A workload dominated by repeated directory removal should not be assumed to scale like random sample reads. The paper explicitly observes that removal overhead grows with the metadata-node population.

Correctness also depends on ordering a local operation against invalidation. If an operation already holds the relevant directory lock, invalidation waits; if invalidation wins, a subsequent lookup must obtain valid information before proceeding. This is not an eventual-consistency shortcut that lets an unauthorized operation succeed temporarily. The fast path comes from relocating checks and sharing metadata, not abandoning them.

The Linux interface and its compatibility boundary

Linux’s virtual filesystem layer already performs path walking and caches directory objects. Simply moving resolution to the server would leave that client work in place. FalconFS’s client distinguishes intermediate path components from the final target and uses temporary attributes to let the intermediate local walk proceed. The server performs the real existence and permission checks for the complete path.

Those temporary attributes must not become user-visible file properties. The client revalidates entries when they become the final component of an operation and fetches real attributes as needed. This is a subtle implementation requirement, not a suggestion that other filesystems should return permissive attributes without equivalent server checks and revalidation.

The paper also states concrete limitations: symbolic links and nested mount-point handling are unsupported in this implementation. Directory access and modification timestamps are not maintained in the same way as the comparison systems. These differences can matter to dataset tools, backup jobs, packaging scripts, and applications that use timestamp changes to detect updates.

A POSIX-like interface is therefore not sufficient evidence for a transparent migration. Before deployment, enumerate the operations used by the entire pipeline, including ingestion, labeling, archiving, and administration. A training loop may use only ordinary reads while the scripts around it depend on precisely the semantics that differ.

Batching trades individual latency for shared work

Once requests meet at a metadata server, FalconFS can merge concurrent work. Requests with common directory prefixes can share parts of the locking effort. Metadata changes can also amortize write-ahead-log costs across a batch instead of issuing many small persistence operations independently. PostgreSQL supplies underlying transaction and logging machinery, with filesystem-specific extensions around it.

This is a throughput optimization, not a free reduction in every request’s latency. Waiting for or processing a batch can increase individual operation time. The paper reports cases where Lustre has lower latency or better throughput at lower concurrency, while FalconFS benefits as concurrency grows. The workload must supply enough simultaneous work for amortization to matter.

That distinction suggests separate tests for a large training fleet and an interactive user browsing a dataset. They may access the same files but have very different concurrency and response-time requirements. A single peak operations-per-second result cannot describe both experiences.

What the training experiment actually measures

The experimental storage cluster is built from 13 dual-socket machines, divided into 26 resource-bound logical nodes. Metadata-server CPU resources are capped for saturation tests. The comparisons include specific versions of CephFS, JuiceFS, and Lustre, and disable both metadata and data replication. Some peak metadata measurements use a library client because the available FUSE clients cannot saturate the servers.

These details prevent an immediate translation into a production purchasing ratio. Replication changes persistence and network costs; client interfaces change CPU overhead; software versions and configuration affect the baseline. The experiments isolate architectural effects, but an operator needs a second round with its intended durability settings and client path.

For training, the authors use MLPerf Storage to emulate a ResNet-50 input workload. It reads ten million 112 KiB files across one million directories with direct I/O. At a 90% accelerator-utilization threshold, the reported supported emulated accelerator counts are 80 for FalconFS and 32 for Lustre. CephFS does not meet that threshold in this test. This is not a physical 80-GPU training run and does not establish zero useful capacity for CephFS in other settings.

Supported emulated accelerator counts at the paper’s 90% utilization threshold: FalconFS 80 and Lustre 32. The workload is an MLPerf Storage ResNet-50 simulation with direct I/O, and benchmark replication is disabled. CephFS did not meet the threshold and is not plotted as zero capacity. Original figure created for this article.

The final PDF’s abstract and evaluation text do not consistently state the same maximum speedup. We therefore do not use that maximum as the article’s headline result. The explicitly described threshold comparison, request counts, and mechanism are more useful for a design decision than choosing the larger number from inconsistent passages.

Production use is not a substitute for a failure test

The paper separately reports a year of production use in Huawei’s autonomous-driving environment with 10,000 NPUs. That is evidence that the system served a substantial operating pipeline. It should not be merged with the replication-disabled experimental configuration as if the entire production fleet produced the benchmark results.

FalconFS describes logging, primary-secondary replication, and coordinator recovery. These mechanisms are distinct from the lazy directory replicas used to accelerate path resolution. Confusing the two would make a namespace optimization appear to be the complete durability design. Metadata availability, committed-write recovery, and file-data protection each need their own failure model.

Resizing is another explicit limitation. In the described implementation, moving inodes during cluster reconfiguration stops request service rather than providing live migration. A storage system selected for an expanding AI platform must therefore be tested against the expected resize frequency and maintenance window. High steady-state throughput cannot compensate for an unacceptable interruption during capacity growth.

Choosing between a filesystem change and a dataset change

Combining small samples into larger containers can also reduce file-open work. That alternative changes the data layout or loader rather than the filesystem. It may be attractive for a stable training dataset, while independently writable samples and multiple pipeline stages can make such repacking less convenient. Neither approach should be declared universally superior without examining who reads and modifies the data.

The useful first measurement is the number of metadata requests per useful sample, together with CPU and memory consumed by the client. If those costs dominate, FalconFS offers a concrete architecture to investigate. If decompression, augmentation, SSD bandwidth, or networking dominates instead, improving path resolution may leave end-to-end progress nearly unchanged.

The final decision should be based on supported accelerators at the required utilization and durability, not raw SSD bandwidth or one maximum speedup. Repeat the workload with the real directory distribution, naming conventions, permissions, concurrency, and growth procedures. FalconFS is most compelling when a large shared namespace makes client-side repetition expensive and the application can accept its clearly stated operational constraints.

This independent editorial digest uses the final NSDI 2026 proceedings paper[1], authored by researchers at Huawei Technologies and Shanghai Jiao Tong University. Measurements and implementation limits are attributed to that version; deployment recommendations are our analysis. No source table or figure was reproduced. The text and explanatory graphics were created for Silicon & Systems. Original-paper copyright remains with its rights holders, © 2026; USENIX Association publishes the proceedings.