siliconandsystems.com/en/articles

Articles

08-31[news]The Memory Controller Changes Sides: What NVHBM Really Buys21m08-25[news]Three Wafers Become One Service Unit: Reading Cerebras CS-4 Beyond Token Speed12m08-25[memory]Mobile Memory Takes the Server Socket: LPDDR6 Silicon, SOCAMM2 and DRAM That Computes19m08-25[arch]Not a Smaller GPU: How Jalapeño Rebuilds Inference Around Locality25m08-24[news]The 72-GPU Rack Is the Product: How MI455X Becomes AMD Helios12m08-24[news]One Core, Two Instruction Worlds: Why IBM Is Bringing Arm Inside the Mainframe12m08-24[memory]The Boundary Learned to Move: 52 Industry Memory and Storage Papers in 202628m08-24[news]Three Chips, Three Power Budgets: Intel’s Agentic AI Portfolio Test12m08-24[arch]Vera Rebalances the CPU: 88 Cores Fed by 1.2 TB/s19m08-17[fabric]One DPU, Eight RNICs: Tencent's Pegasus Network for Bare-Metal AI Cloud15m08-10[fabric]Beyond NVLink and Scale-Out Networks: A One-Chip-Like CXL Datacenter33m08-06[fabric]A Flow-Size Distribution Cannot Describe a Burst20m08-03[memory]NAND Moves Up a Tier: Reading the First High Bandwidth Flash Specification20m07-26[arch]RTL Automation Starts After the First Draft15m07-15[silicon]Arrays, Timing, and Structured Compute: Industry Circuits in 202618m07-13[ai-infra]The Datacenter Became the Runtime: Nine Industry Papers That Redefined AI Infrastructure23m07-13[snapshot]Batch Analytics Is Idle Even When Its CPU Is Allocated14m07-13[ai-infra]Idle Is Not Available: The Accounting Problem Inside Alibaba's 155,410-GPU Fleet19m07-13[snapshot]A Flat Profile Can Make a Faster Program Look Slower12m07-13[snapshot]Serve Linearizable Reads from the Node the Client Needs14m07-13[snapshot]An Exabyte Data Lake Can Feed LLM Training If It Learns the Schedule19m07-13[snapshot]Unify Logs, Metrics, and Traces Without an In-Memory Fleet13m07-13[snapshot]Long-Lived VMs Leave Placement Debt Behind14m07-13[snapshot]RL Training Needs a Scheduler Inside the Job12m07-13[snapshot]An LLM Optimization Is a Production Change, Not a Suggestion12m07-13[ai-infra]A Silent GPU Error Needs Three Different Questions19m07-13[snapshot]Put Billion-Vector Search Back on Flash14m07-13[snapshot]A Rack Power Budget Changes with Hardware Age14m07-13[snapshot]Instruction-Level iOS Analysis Without Generating Executable Code13m07-13[snapshot]Put the VM Observer Beside the VMM14m07-13[snapshot]Split CPU and Memory Virtualization Across Worlds14m07-13[snapshot]Verify the Allocator That Runs on Twelve Million Phones14m07-13[snapshot]Separate Log Durability from Log Order14m07-13[snapshot]Datacenter Memory Reclamation Reverses the Cache Problem14m07-13[snapshot]Follow the User Interaction Across Android Processes14m07-13[ai-infra]The Agent Graph Becomes the Cloud Scheduler: Murakkab Optimizes the Whole Workflow19m07-13[snapshot]Memory Tiering Cannot Reclaim a Cold Object Trapped on a Hot Page12m07-13[snapshot]One Fault-Domain Buffer Can Carry Fleet Maintenance14m07-13[snapshot]Clear Accelerator Demand Every Minute, Not Every Quarter13m07-13[memory]Memory Capacity Is Not Memory Bandwidth: RamRyder Gives Cloud VMs Their Own Channels17m07-13[snapshot]Replay the GPU Step That Corrupted the Model14m07-13[snapshot]Learn the Rollout Tail Before It Forms14m07-13[snapshot]Make the S3 Model Precise Enough to Be the Specification13m07-13[memory]A Cache Hit That Arrives Late Is Still a Miss: Strata Schedules Long-Context Memory19m07-13[ai-infra]Trace the Request That Missed Its SLO, Not Every Request Around It19m07-13[ai-infra]The Pipeline Is No Longer Made of Identical Bricks19m07-13[ai-infra]Move the Training Job Before It Learns That a Machine Left18m07-13[ai-infra]Fill the Idle Half of On-Policy RL: Weave Co-Schedules Rollout and Training20m07-13[snapshot]Turn Kernel Constants into Millisecond Control Points12m06-27[silicon]AmpereOne Hides Memory Tags in the Bits a Server Already Pays For19m06-27[silicon]Li Auto Lets the Compiler Drive Data Instead of Rebuilding a GPU Cache20m06-27[silicon]The PCIe Tax on CXL, and the 4 nm Silicon Built to Repeal It22m06-27[ai-infra]The Weather File Expires Before the Datacenter: Google Reprices Cooling for 204424m06-27[silicon]Dropping the PHY: d-Matrix Bonds Compute Directly onto DRAM19m06-27[silicon]Microsoft's Rowhammer Defense Sleeps Until One Sub-Bank Looks Dangerous20m06-27[memory]CXL Leaves the Lab: How Meta Turned Retired DDR4 into a Production Memory Tier21m06-16[silicon]Three Nanosheets over Three: Samsung Builds Logic Upward at 42 nm14m06-01[arch]Production Evidence Replaced the Architecture Promise: 43 Company-Affiliated Papers at ISCA 202616m05-23[ai-infra]Power Is the Cluster: What 83,000 GB200s Teach About 150 MW28m05-19[ai-infra]Reasoning Models Hit the Memory-Capacity Wall Before the Compute Ceiling21m05-18[ai-infra]Millions of Valid Configurations, One Deployment That Meets the SLO18m05-04[ai-infra]Fast Tokens, Slow Agents: Agentix Schedules the Program Behind Each LLM Call16m05-04[fabric]Keeping Old NICs on the RDMA Path: What ByteDance's BURST Really Accelerates17m05-04[ai-infra]Balanced Load, Unbalanced Latency: Alibaba Protects Quiet Storage Work from Bursts16m05-04[ai-infra]The Same Context, Different Weights: DroidSpeak Reuses KV State Between Fine-Tuned Models16m05-04[ai-infra]The GPU Is Busy, but Training Is Slow: EROICA Finds the Faulty Step Online13m05-04[memory]Before the GPU Reads a Byte: FalconFS Moves the Directory Walk Off Training Clients17m05-04[ai-infra]Inference Owns the Deadline, Fine-Tuning Uses the Gaps: FlexLLM Schedules Tokens Together13m05-04[fabric]A 100K-GPU Network Starts as a Compiler Problem22m05-04[fabric]Compile the Fabric before Deploying It: Meta's Matryoshka Network Design System15m05-04[memory]No Giant Switch: How Octopus Builds a 96-Server CXL Memory Pod13m05-04[fabric]Azure Puts Stateful Networking Inside the Switch: The Deployment Logic of SONiC DASH18m05-02[arch]SPEC CPU 2026 Stops Pretending Every Core Runs the Same Program19m04-26[ai-infra]One Batch, Different Deadlines: AdaServe Builds a Speculation Tree for Each SLO12m04-26[ai-infra]The Bits That Never Change: IBP Compresses the PCIe Path Without Changing the Model12m04-26[fabric]The Fast Link Is Not the Only Link: MPCCS Stripes GPU Collectives Across Two Fabrics12m04-22[ai-infra]One TPU Generation, Two Networks: Why Google Split Training from Reasoning21m03-24[news]Arm Stops at the Rack: What Its First Production CPU Changes12m03-24[memory]A Real CXL Memory Box Under SAP HANA: What Can Move, and What Must Stay Local13m03-22[ai-infra]The Same H100 Is Not the Same Rental: Measuring the GPU Cloud Lottery18m03-22[arch]AI Systems Turned Every Layer into a Scheduling Problem: 44 Company-Affiliated Papers at ASPLOS 202615m02-24[memory]Five Regions Changed the Economics of Apple's Object Store16m02-24[ai-infra]When Storage Understands the Job: AITURBO Turns AI I/O into a Group Operation19m02-24[snapshot]Cloud Garbage Collection Can Punch Holes Before It Moves Bytes13m02-24[memory]KV Reuse Breaks When the Text Moves, Even If the Text Does Not17m02-24[memory]The SSD Stayed Local, but Its Control Path Moved Three Times17m02-24[snapshot]Container Startup Is a Metadata Path Before It Is a Download13m02-24[snapshot]An SSD Failure Alert Must Explain the Combination That Fired12m02-24[snapshot]Data Integrity Must Cross the Filesystem Boundary14m02-24[snapshot]Database Compression Needs a Fast Path for the Bytes That Commit13m02-24[memory]The Fastest Model Loader Changed the Page Cache, Not the Framework16m02-24[snapshot]File Synchronization Can Reuse Work the Storage Layer Already Did13m02-24[memory]Compression Improves When the Filesystem Sorts Before It Packs14m02-24[memory]A Thousand Cartridges, Four Drives: Why Tape Cloud Storage Must Wait on Purpose19m02-24[snapshot]Zoned Mobile Flash Works Only When Android Preserves Its Contract13m02-17[silicon]What Commercial AI EDA Has Actually Put Through Tapeout21m02-17[fabric]The 400 µm Modulator That Pushes CPO to 212 Gb/s per Wavelength19m02-10[silicon]The Next EDA Advantage Is a Recipe That Transfers15m01-31[arch]Scale Exposed the Hidden Architecture Bill: 28 Company-Affiliated Papers at HPCA 202613m01-30[fabric]Who Wires the Rack: Reading UALink Against NVLink and Ethernet21m01-26[news]Ethernet Inside the Scale-Up Domain: The System Bet Behind Maia 20012m01-21[silicon]Place and Route Is Becoming an Objective Negotiation14m
# 2025
12-29[ai-infra]Meta Gives Kernel Optimization a Search Tree and a Memory20m12-16[ai-infra]What a Neocloud Number Proves, and What It Leaves Unmeasured20m12-08[silicon]CFET Leaves the Device Lab: TSMC Runs an Oscillator and SRAM below 48 nm14m11-25[memory]KVCache Without the Network: Alibaba Runs LLM Serving on a Real CXL 2.0 Switch19m11-20[arch]AI Can Search a Microarchitecture, but It Still Needs an Executable Contract15m11-19[silicon]Evatec Uses Film Stress to Flatten the Packaging Wafer16m11-03[memory]The Hierarchy Became the Product: 34 Industry Memory and Storage Papers from 202529m10-23[fabric]One RoCE Fabric, Four Distance Classes: Meta's Communication Stack for 100K+ GPUs15m10-17[arch]Implementation Details Became System Constraints: 35 Company-Affiliated Papers at MICRO 202514m10-13[ai-infra]Seven Models, One GPU: What the Long Tail Costs and How Aegaeon Reprices It20m10-13[ai-infra]Evict First, Ask Questions Later: How ByteDance Keeps 9,600 GPUs Worth Training On19m10-12[snapshot]Memory-Safe Compartments Without an MMU14m10-12[snapshot]Reuse Yesterday's Optimizer State for the Next Allocation Round14m10-12[snapshot]Compress Keys, Values, Tokens, and Heads Differently14m10-12[ai-infra]When Remote HBM Becomes a Compiler-Managed Memory Level18m10-12[ai-infra]Stop Treating the Network as a Barrier: Mercury Compiles Remote HBM into the Operator14m10-12[fabric]Pooling PCIe Devices Without a PCIe Switch14m10-12[snapshot]One Output Token Changes the Inference Engine14m10-12[snapshot]Freeze the RDMA Device, Not the Application Contract14m10-12[snapshot]Let SmartNIC Control Tasks Borrow Data-Plane Cores14m09-28[fabric]One Photonic Switch, Two Routing Dimensions14m09-10[silicon]At 100 GHz, the Laser Package Becomes Part of the Circuit14m09-08[fabric]An AI Fabric Needs Three Views of the Same Failure19m09-08[fabric]The FLOPS You Cannot Buy, You Wire: Inside Tencent's Half-Million-GPU Fabric20m09-08[fabric]Ask the Switch Which Path the Probe Actually Took18m09-08[silicon]Precision Moves Into the Architecture: Industry Circuits in 202519m09-08[fabric]Firing the Switch: A Scale-Up Domain Built from Transceivers19m09-08[fabric]A Local Queue Cannot See the Congestion Two Switches Away18m09-08[fabric]The Training Job Already Knows Which Network Paths Matter18m08-27[ai-infra]Alibaba Rebuilt Cloud RDMA Across Host, PCIe, and Fabric15m08-27[silicon]The Soft Materials Set the Shape of the Advanced Package16m08-26[fabric]UCIe Crosses the Rack in Light: Ayar Labs' 8.192 Tb/s Optical Retimer15m08-26[fabric]Optics Moves Inside the Silicon: Celestial AI's Photonic Fabric Module15m08-26[fabric]A 4,000 mm² Optical Backplane: Reading Lightmatter's Passage M100016m08-06[arch]The Hardware-Software Contract Moved Again: 59 Company-Affiliated Papers at ASPLOS 202516m08-06[fabric]An RDMA NIC Needs a Scheduler, Not Just Faster Queues14m08-04[ai-infra]When Thousands of GPUs Wait for Storage: Reading MLPerf Storage 2.019m07-07[snapshot]One Publishing Path for Live and Collaborative Video13m07-07[ai-infra]A New GPU Can Help Before the Model Finishes Loading: BlitzScale's Live Autoscaling17m07-07[snapshot]When One Ordered Stream Must Reach Thousands of Readers16m07-07[snapshot]A Failure Detector Fast Enough to Change the Protocol13m07-07[snapshot]Recover the Reduction Tree from Floating-Point Outputs13m07-07[ai-infra]The GPU Did Not Crash, Yet 1,024 Workers Slowed Down19m07-07[ai-infra]Inside the GPU Kernel: KPerfIR Makes Compiler Decisions Measurable16m07-07[snapshot]Record the Kernel Slice, Not the Entire Machine14m07-07[snapshot]Spend Latency When the Shopper Stops Moving13m07-07[ai-infra]Why Elastic Tensors Help PCIe GPU Servers More13m07-07[snapshot]The Receive Ring Was Doing Two Jobs12m07-07[memory]Read Layout Is Not a Recovery Group: Okapi Separates Two Storage Decisions18m07-07[snapshot]Move Paging Policy Out, Keep the Fault Path In13m07-07[snapshot]One Training Contract Across Ten Million CPU Cores12m07-07[snapshot]When an LLM Writes the Port and a Solver Repairs It13m07-07[snapshot]A Sub-Millisecond Fork Does Not Make a Fast Function14m07-07[ai-infra]Inference Can Borrow the Whole GPU: SIRIUS Makes Training Return Memory in Milliseconds15m07-07[snapshot]A Two-Second Freshness Contract Without Rebuilding the Cache15m07-07[ai-infra]When Idle Models Leave the GPU: Paying Only for Active Inference19m07-07[ai-infra]A Million Cores Do Not Behave Like One GPU: WaferLLM Rewrites Inference for the Mesh16m07-07[snapshot]Long Contexts Turn Equal Tokens into Unequal Work14m07-07[arch]One Scheduler for Ten Accelerators: XSched's Three Levels of Preemption19m06-22[silicon]Why AI Reached Industrial Physical Design Before Architecture15m06-21[arch]A CPU Benchmark That Meta Was Willing to Buy Servers With19m06-21[ai-infra]The Hardware Wishlist DeepSeek Wrote on a Halved NVLink21m06-21[ai-infra]Why Llama 3 Needed Four Kinds of Parallelism at Once20m06-21[silicon]No HBM, No Apologies: MTIA 2i and Meta's Productionization Ledger19m06-20[arch]Architecture Reached the Fleet: 43 Company-Affiliated Papers at ISCA 202515m05-28[fabric]CPO Has to Survive Reflow: Intel's Fiber-to-EMIB Package18m05-27[silicon]IBM Moves the Heat Spreader into the Organic Substrate16m05-14[fabric]Pool the Buffer, Not the Device: A CXL Shortcut Around PCIe Switches13m05-12[ai-infra]Long Contexts Leave GPU Waves Half Empty: LeanAttention Rebalances Decode15m05-01[memory]Eighty Percent Far, Eighty-Five Percent Fast: A Database Placement Rule for CXL12m04-28[ai-infra]A Checkpoint Should Outlive the GPU Layout: ByteCheckpoint Saves Logical Tensors12m04-28[memory]RDMA Remote Memory Without a Memory-Node CPU15m04-10[ai-infra]Three Percent Globally, a Grid Problem Locally: Reading the IEA's AI Forecast20m03-30[ai-infra]Training Already Sends the Gradients: FlowCheck Rebuilds Checkpoints from Mirrored Traffic13m03-30[ai-infra]One Controller Above, Many Below: HybridFlow, the Engine Inside verl19m03-30[arch]Prefill Has Compute, Decode Has Bandwidth: POD-Attention Puts Both on the Same SM13m03-30[ai-infra]Cloud Isolation Without Leaving the Fast Path: What Vela Learned at 1,500 GPUs12m03-01[ai-infra]Alibaba Makes Collective Communication Diagnose the Cluster21m03-01[arch]Measured Architecture Became an Operating Policy: 40 Company-Affiliated Papers at HPCA 202515m03-01[ai-infra]The Largest GPU Jobs Fail Most, but Small Jobs Still Set Fleet Policy21m03-01[ai-infra]AI Energy Needs a Meter That Survives Nine Orders of Magnitude21m02-25[snapshot]A Filesystem Process Can Crash Without Erasing the Application's Past14m02-25[snapshot]A TEE Must Reject Disk States the Application Never Committed13m02-25[snapshot]Container Startup Wants a Memory Image, Not a Storage Archive13m02-25[ai-infra]Buying Back FLOPs with DRAM: Mooncake, the Cache That Serves Kimi21m02-17[silicon]A 600 mm² Mainframe Die Becomes a System: Inside IBM Telum II16m02-16[memory]HBM Saves Clock Power by Rebuilding Four Phases at the Data Groups24m01-30[fabric]Let the Optical Topology Follow the GPU Allocation14m01-23[fabric]Count the Laser: Co-Designing a 224 Gb/s Coherent CPO Link18m
# 2024
12-11[silicon]When AI Placement Becomes Production Silicon21m09-09[silicon]What Industry Put on Silicon in 202418m08-06[memory]Memory Stopped Being a Component: 35 Industry Papers from 202429m06-29[silicon]MI300A Turns the Accelerator Package into the Compute Node19m06-29[silicon]FuriosaAI Replaces the Fixed Matrix with a Shape-Shifting Tensor Engine20m04-16[fabric]Routable PCIe Is a Fabric, Not a Longer Bus12m04-14[fabric]Spend Delay Before the First Hop to Buy Back Bandwidth21m
# 2023
03-24[arch]Making a SmartNIC Fast Across a Slow PCIe Boundary13m02-07[fabric]The Protocol Debt Inside Hyperscale RoCE14m