An AI training job and a network operator describe the same incident differently. The job reports a slow or stuck collective. The operator sees links, switch ports, overlay tunnels, adaptive paths and queues. One failed component may affect only a container placement, while one congested path may be healthy in the binary sense but still destroy collective goodput. A useful operating stack must answer three questions: which part of the expected training graph stopped behaving, which physical path carried the affected packets, and where traffic should move next.
Three industry-led SIGCOMM 2025 papers cover those control points. Alibaba Cloud’s SkeletonHunter infers failures from the regular, sparse traffic of large-model training[1]. ByteDance’s ByteTracker sends centralized probes and uses switch mirroring to recover their paths without installing agents on every server[2]. ByteDance and Broadcom’s SGLB gives commodity switches a global congestion-aware load-balancing mechanism[3]. They should not be ranked by one metric. The first two diagnose production networks; the third evaluates a forwarding prototype that acts before or during congestion.
The workload already draws a skeleton
Containers make placement dynamic and add an overlay network over the physical fabric. Comprehensive monitoring of every possible path is expensive, while opportunistic samples may miss the paths that matter to the current job. SkeletonHunter uses a property of distributed training: its communication is large but structurally sparse and repetitive. Tensor, pipeline, data and expert parallel groups repeatedly traverse a limited set of paths. This recurring set is the traffic skeleton.
The system identifies the expected skeleton and reasons about deviations in its signals. Because the selected paths are tied to active training communication, a missing or degraded observation has stronger workload context than a generic background probe. The overlay-to-underlay relationship also helps translate a container-visible symptom into candidate physical components.
Alibaba reports six months of production deployment. SkeletonHunter uncovered 4,816 network failures with 98.2% precision and 99.3% recall, and localized them with 95.7% accuracy. After operators fixed 98% of the identified problematic components, the monthly network failure rate fell by 99.1%. The last number is an operational before-after result, not an algorithmic localization score. Maintenance action sits between detection and the reduction.
Recover a path without managing an agent fleet
ByteTracker begins from a different gap. End-host probes add processes, configuration and host noise across a data center that may contain an enormous number of servers. A timeout can come from the host rather than the network, and adaptive forwarding makes the actual path harder to infer from endpoints alone.
The system places a small number of dedicated probers centrally. It constructs packets that cause the destination operating system to return a response without a resident probing agent. North-south and east-west probe patterns cover links and internal forwarding paths. Switch packet mirroring records the actual route and enables hop-by-hop comparison, including anomalies such as silent drops or changed packet contents.
ByteTracker was deployed in all ByteDance data centers for more than half a year. The official paper abstract reports that it detected almost all network anomalies and localized them within five seconds with 100% accuracy during the deployment. That result should retain its context: it is the authors’ operational evaluation over the anomalies observed in that period, not a mathematical guarantee for every failure mode. Its advantage over SkeletonHunter is not a higher percentage. It covers general data-center paths without relying on an active training pattern; SkeletonHunter adds training-topology semantics.

Steer around congestion before it becomes a failure
SGLB acts in the forwarding plane rather than the incident queue. Local adaptive routing sees nearby queues but can push traffic toward congestion farther along a path. Global state is more informative but expensive to distribute and store in commodity switch hardware. Link failure also demands fast convergence, while asymmetric path capacity can make naïve traffic spreading reduce throughput.
The system’s SyncMesh control protocol distributes compact congestion profiles to a Global Load Balancing engine in switches. The design combines global path condition with available capacity and updates forwarding after failures. Prototype experiments report recovery in as little as 45 microseconds and up to 60% faster All-to-All communication by avoiding globally congested paths.
These are not six-month production figures like those for SkeletonHunter and ByteTracker. The paper prototypes SGLB and evaluates it experimentally. It establishes that the abstraction can fit commodity switch constraints and improve selected collectives under tested conditions. Long-term behavior under changing job mixes, control-plane loss and interacting traffic remains a deployment question.

A combined operating loop
The systems fit together as a loop. Training traffic supplies high-value passive evidence. Dedicated probes test paths independently of the workload and retain coverage during job changes or outages. The load balancer uses current congestion state to avoid damage, while both monitoring systems confirm whether steering restored expected behavior. When a component is truly bad, localization feeds maintenance and quarantine rather than endless rerouting.
There are seams to engineer. Mirrored probes must follow paths representative of GPU traffic. Adaptive routing can make one probe observation stale quickly. A traffic skeleton changes when the scheduler moves containers or parallel groups. Global congestion summaries have delay and finite precision. The operator therefore needs common identifiers across job, container, endpoint, switch and path, plus timestamps accurate enough to join their evidence.
The larger lesson is that “network reliability” is not one percentage. Detection precision, recall, localization accuracy, time to localization, convergence time and collective throughput belong to different stages. An AI cloud that publishes one of them has not automatically demonstrated the others. The industry papers are valuable precisely because they expose several stages of the operating system behind the fabric.
The integration problem is larger than any one result
The three papers should not be read as competing monitoring products. They observe different clocks. SkeletonHunter follows the stable path structure of a long-running job and detects a change. ByteTracker resolves the physical path of active probes within seconds. SGLB reacts inside the network on a microsecond scale before a control-room diagnosis could arrive. A complete fabric needs all three clocks because detecting, explaining, and steering around a fault are different deadlines.
Their metrics also belong to different denominators. Precision and recall describe a classifier over observed failures. Five seconds describes operational localization latency across a deployed estate. Forty-five microseconds and 60% describe recovery and collective performance in a prototype experiment. The useful synthesis is not a ranking but an interface: the workload layer should identify the affected skeleton, the path layer should map it to components, and the control layer should keep traffic productive while repair proceeds. Without that handoff, each tool can be locally correct while the job still stalls.
This suggests a stronger acceptance test for an AI fabric. Measure completed training steps during injected link, switch, and congestion events, then decompose the loss into detection delay, localization delay, route convergence, degraded-path bandwidth, and repair time. Include false alarms that trigger unnecessary rerouting, because a control loop that reacts quickly to the wrong signal can create the outage it was meant to avoid. The headline should be the fraction of target collective goodput preserved through the event, not merely the time when an alarm appeared. The larger insight is that network operations are becoming part of the distributed runtime. Job topology supplies context to telemetry, telemetry supplies evidence to routing, and routing buys time for physical repair.
The workload skeleton also changes what deserves monitoring. A conventional network tool tries to cover every path uniformly. Large distributed jobs repeatedly exercise a sparse and structured subset, so the marginal value of observing those paths is higher. However, the subset changes when a job is rescheduled, a collective algorithm changes, or elastic membership moves ranks. The monitoring system must version the skeleton with the job plan; otherwise it can confidently watch yesterday’s critical paths while today’s traffic fails elsewhere.
Control-loop stability is the corresponding SGLB concern. Global congestion information is more powerful than a local queue sample, but it is older by the time it reaches the switch. If many switches react to the same delayed summary, they can move traffic together and create a new hot path. Prototype recovery time does not by itself establish stability under many overlapping jobs and failures. Production evidence should include oscillation, fairness, and convergence under stale or partial state. The operator’s goal is not globally perfect information. It is information fresh and bounded enough that distributed decisions improve the next interval instead of correcting the last one.
The three views need a common evidence model
Workload structure, physical path, and congestion state are produced by different systems and clocks. A job scheduler knows ranks and collectives. Network telemetry knows ports, queues, and links. Tracing knows when a request or step slowed. Without a shared identifier and time model, operators can place three correct dashboards beside one another and still fail to explain the same event.
The evidence model should begin with a job, communication group, collective or request phase, and the set of expected endpoints. Path reconstruction can attach the actual devices and links used during that interval. Congestion telemetry can then attach queue occupancy, drops, pause events, or reroutes to those links. Each attachment needs a confidence and freshness field because sampled telemetry, inferred paths, and scheduler intent are not equally certain.
Clock alignment is part of correctness. A congestion spike observed after a collective completed is not its cause, even if both occur on the same port. Controllers and hosts should synchronize time closely enough for the shortest diagnosed event, or carry causal sequence markers when wall-clock accuracy is insufficient. The incident view should preserve raw intervals so an automated correlation does not turn temporal proximity into an unsupported root cause.
Control loops must have explicit ownership
Several layers can react to the same signal. The transport may reroute traffic, the network controller may alter paths, and the job scheduler may migrate or pause work. Independent reactions can oscillate: a job moves away from congestion just as the network repairs it, creating a new hotspot elsewhere. A combined operating loop therefore needs authority boundaries and hold-down times.
Fast local mechanisms should handle brief faults they can resolve without changing job placement. The network controller can manage persistent path pressure within the fabric. The scheduler should act when the expected remaining cost of a bad placement exceeds migration, checkpoint, or restart cost. Each layer must report the action it took and the interval during which upper layers should observe rather than respond.
Fallback behavior matters when telemetry or control is unavailable. The network should continue forwarding with conservative defaults, and the scheduler should not infer health from missing data. A stale path map must expire. Any automated remediation needs a bounded blast radius, a rollback condition, and a way to disable one layer without losing all observability.
Operations should optimize completed synchronized work
Network utilization is not the final objective for AI clusters. A lightly used link can sit on the critical path, while a heavily used link can carry traffic fully hidden by compute. The service metric should connect network events to completed training steps, inference requests, or checkpoint intervals and their SLOs.
For training, useful reporting includes step-time distribution, exposed collective time, straggler ranks, and lost accelerator-hours. For inference, it includes TTFT or token latency by path and queue state. The denominator should identify the affected job or tenant. A fleet-wide average can make a severe but localized regression disappear.
The operating review should also record false interventions. Rerouting or migration that does not improve the critical path consumes bandwidth and destabilizes placement. Comparing predicted improvement with observed outcome builds a calibration set for the control policy. Successful fixes become regression scenarios; failed fixes become constraints that prevent the same reaction.
A staged deployment reduces correlated mistakes
The first stage joins data without taking action. Operators validate endpoint identity, path reconstruction, clock alignment, and the association between congestion and workload delay. The second stage issues recommendations and compares them with human decisions. The third permits reversible local actions with narrow scope. Cross-layer migration or admission changes come last because they can affect many jobs.
Each stage should include injected failures: a disabled link, a congested queue, a misrouted rank, a silent telemetry gap, and a controller restart. The goal is not only to detect the incident but to show that stale or contradictory evidence does not trigger a larger failure. This is particularly important because the combined system has more authority than any one of its source papers.
We read the three works together as an argument for causal operations. Workload knowledge tells the network which paths matter, path knowledge tells telemetry where to look, and congestion state tells the scheduler whether moving work can help. The value appears only when those statements refer to the same event and one control layer owns the response. The product is not another dashboard. It is fewer synchronized jobs losing time to network conditions that the infrastructure could have seen and repaired.
Sources and attribution
This article is an independent synthesis prepared by Silicon & Systems from three industry-led SIGCOMM 2025 papers. Their arguments and results are restated in our own words. No source text, tables or figures are reproduced; both figures were created for this article. Paper metadata and abstracts are available from the official SIGCOMM 2025 program. Copyright (c) 2025 the respective authors and publishers.