siliconandsystems.com/en/articles/category/ai-infra

AI Infrastructure

07-13[ai-infra]The Datacenter Became the Runtime: Nine Industry Papers That Redefined AI Infrastructure23m07-13[ai-infra]Idle Is Not Available: The Accounting Problem Inside Alibaba's 155,410-GPU Fleet19m07-13[ai-infra]A Silent GPU Error Needs Three Different Questions19m07-13[ai-infra]The Agent Graph Becomes the Cloud Scheduler: Murakkab Optimizes the Whole Workflow19m07-13[ai-infra]Trace the Request That Missed Its SLO, Not Every Request Around It19m07-13[ai-infra]The Pipeline Is No Longer Made of Identical Bricks19m07-13[ai-infra]Move the Training Job Before It Learns That a Machine Left18m07-13[ai-infra]Fill the Idle Half of On-Policy RL: Weave Co-Schedules Rollout and Training20m06-27[ai-infra]The Weather File Expires Before the Datacenter: Google Reprices Cooling for 204424m05-23[ai-infra]Power Is the Cluster: What 83,000 GB200s Teach About 150 MW28m05-19[ai-infra]Reasoning Models Hit the Memory-Capacity Wall Before the Compute Ceiling21m05-18[ai-infra]Millions of Valid Configurations, One Deployment That Meets the SLO18m05-04[ai-infra]Fast Tokens, Slow Agents: Agentix Schedules the Program Behind Each LLM Call16m05-04[ai-infra]Balanced Load, Unbalanced Latency: Alibaba Protects Quiet Storage Work from Bursts16m05-04[ai-infra]The Same Context, Different Weights: DroidSpeak Reuses KV State Between Fine-Tuned Models16m05-04[ai-infra]The GPU Is Busy, but Training Is Slow: EROICA Finds the Faulty Step Online13m05-04[ai-infra]Inference Owns the Deadline, Fine-Tuning Uses the Gaps: FlexLLM Schedules Tokens Together13m04-26[ai-infra]One Batch, Different Deadlines: AdaServe Builds a Speculation Tree for Each SLO12m04-26[ai-infra]The Bits That Never Change: IBP Compresses the PCIe Path Without Changing the Model12m04-22[ai-infra]One TPU Generation, Two Networks: Why Google Split Training from Reasoning21m03-22[ai-infra]The Same H100 Is Not the Same Rental: Measuring the GPU Cloud Lottery18m02-24[ai-infra]When Storage Understands the Job: AITURBO Turns AI I/O into a Group Operation19m
# 2025
12-29[ai-infra]Meta Gives Kernel Optimization a Search Tree and a Memory20m12-16[ai-infra]What a Neocloud Number Proves, and What It Leaves Unmeasured20m10-13[ai-infra]Seven Models, One GPU: What the Long Tail Costs and How Aegaeon Reprices It20m10-13[ai-infra]Evict First, Ask Questions Later: How ByteDance Keeps 9,600 GPUs Worth Training On19m10-12[ai-infra]When Remote HBM Becomes a Compiler-Managed Memory Level18m10-12[ai-infra]Stop Treating the Network as a Barrier: Mercury Compiles Remote HBM into the Operator14m08-27[ai-infra]Alibaba Rebuilt Cloud RDMA Across Host, PCIe, and Fabric15m08-04[ai-infra]When Thousands of GPUs Wait for Storage: Reading MLPerf Storage 2.019m07-07[ai-infra]A New GPU Can Help Before the Model Finishes Loading: BlitzScale's Live Autoscaling17m07-07[ai-infra]The GPU Did Not Crash, Yet 1,024 Workers Slowed Down19m07-07[ai-infra]Inside the GPU Kernel: KPerfIR Makes Compiler Decisions Measurable16m07-07[ai-infra]Why Elastic Tensors Help PCIe GPU Servers More13m07-07[ai-infra]Inference Can Borrow the Whole GPU: SIRIUS Makes Training Return Memory in Milliseconds15m07-07[ai-infra]When Idle Models Leave the GPU: Paying Only for Active Inference19m07-07[ai-infra]A Million Cores Do Not Behave Like One GPU: WaferLLM Rewrites Inference for the Mesh16m06-21[ai-infra]The Hardware Wishlist DeepSeek Wrote on a Halved NVLink21m06-21[ai-infra]Why Llama 3 Needed Four Kinds of Parallelism at Once20m05-12[ai-infra]Long Contexts Leave GPU Waves Half Empty: LeanAttention Rebalances Decode15m04-28[ai-infra]A Checkpoint Should Outlive the GPU Layout: ByteCheckpoint Saves Logical Tensors12m04-10[ai-infra]Three Percent Globally, a Grid Problem Locally: Reading the IEA's AI Forecast20m03-30[ai-infra]Training Already Sends the Gradients: FlowCheck Rebuilds Checkpoints from Mirrored Traffic13m03-30[ai-infra]One Controller Above, Many Below: HybridFlow, the Engine Inside verl19m03-30[ai-infra]Cloud Isolation Without Leaving the Fast Path: What Vela Learned at 1,500 GPUs12m03-01[ai-infra]Alibaba Makes Collective Communication Diagnose the Cluster21m03-01[ai-infra]The Largest GPU Jobs Fail Most, but Small Jobs Still Set Fleet Policy21m03-01[ai-infra]AI Energy Needs a Meter That Survives Nine Orders of Magnitude21m02-25[ai-infra]Buying Back FLOPs with DRAM: Mooncake, the Cache That Serves Kimi21m