ARTICLE ORDER · NEWEST FIRST
Article order
An operator view arranged from the most recently added article.
#001[memory]HBM Saves Clock Power by Rebuilding Four Phases at the Data Groups#002[silicon]What Commercial AI EDA Has Actually Put Through Tapeout#003[silicon]When AI Placement Becomes Production Silicon#004[arch]RTL Automation Starts After the First Draft#005[arch]AI Can Search a Microarchitecture, but It Still Needs an Executable Contract#006[silicon]Why AI Reached Industrial Physical Design Before Architecture#007[silicon]Place and Route Is Becoming an Objective Negotiation#008[silicon]The Next EDA Advantage Is a Recipe That Transfers#009[fabric]Pooling PCIe Devices Without a PCIe Switch#010[ai-infra]Alibaba Rebuilt Cloud RDMA Across Host, PCIe, and Fabric#011[fabric]An RDMA NIC Needs a Scheduler, Not Just Faster Queues#012[ai-infra]Why Elastic Tensors Help PCIe GPU Servers More#013[memory]RDMA Remote Memory Without a Memory-Node CPU#014[fabric]Routable PCIe Is a Fabric, Not a Longer Bus#015[arch]Making a SmartNIC Fast Across a Slow PCIe Boundary#016[fabric]The Protocol Debt Inside Hyperscale RoCE#017[silicon]The Soft Materials Set the Shape of the Advanced Package#018[silicon]Evatec Uses Film Stress to Flatten the Packaging Wafer#019[silicon]IBM Moves the Heat Spreader into the Organic Substrate#020[fabric]Count the Laser: Co-Designing a 224 Gb/s Coherent CPO Link#021[fabric]CPO Has to Survive Reflow: Intel's Fiber-to-EMIB Package#022[fabric]The 400 µm Modulator That Pushes CPO to 212 Gb/s per Wavelength#023[memory]No Giant Switch: How Octopus Builds a 96-Server CXL Memory Pod#024[memory]A Real CXL Memory Box Under SAP HANA: What Can Move, and What Must Stay Local#025[fabric]Pool the Buffer, Not the Device: A CXL Shortcut Around PCIe Switches#026[memory]Eighty Percent Far, Eighty-Five Percent Fast: A Database Placement Rule for CXL#027[fabric]One Photonic Switch, Two Routing Dimensions#028[silicon]At 100 GHz, the Laser Package Becomes Part of the Circuit#029[fabric]Let the Optical Topology Follow the GPU Allocation#030[ai-infra]The GPU Did Not Crash, Yet 1,024 Workers Slowed Down#031[ai-infra]When Idle Models Leave the GPU: Paying Only for Active Inference#032[fabric]One DPU, Eight RNICs: Tencent's Pegasus Network for Bare-Metal AI Cloud#033[silicon]Three Nanosheets over Three: Samsung Builds Logic Upward at 42 nm#034[fabric]Compile the Fabric before Deploying It: Meta's Matryoshka Network Design System#035[silicon]CFET Leaves the Device Lab: TSMC Runs an Oscillator and SRAM below 48 nm#036[fabric]One RoCE Fabric, Four Distance Classes: Meta's Communication Stack for 100K+ GPUs#037[fabric]UCIe Crosses the Rack in Light: Ayar Labs' 8.192 Tb/s Optical Retimer#038[fabric]Optics Moves Inside the Silicon: Celestial AI's Photonic Fabric Module#039[fabric]A 4,000 mm² Optical Backplane: Reading Lightmatter's Passage M1000#040[silicon]A 600 mm² Mainframe Die Becomes a System: Inside IBM Telum II#041[ai-infra]The Same H100 Is Not the Same Rental: Measuring the GPU Cloud Lottery#042[memory]KV Reuse Breaks When the Text Moves, Even If the Text Does Not#043[fabric]A Flow-Size Distribution Cannot Describe a Burst#044[fabric]A 100K-GPU Network Starts as a Compiler Problem#045[fabric]Spend Delay Before the First Hop to Buy Back Bandwidth#046[memory]The Fastest Model Loader Changed the Page Cache, Not the Framework#047[memory]Five Regions Changed the Economics of Apple's Object Store#048[memory]Compression Improves When the Filesystem Sorts Before It Packs#049[memory]The Boundary Learned to Move: 52 Industry Memory and Storage Papers in 2026#050[fabric]Beyond NVLink and InfiniBand: A One-Chip-Like CXL Datacenter#051[silicon]Arrays, Timing, and Structured Compute: Industry Circuits in 2026#052[silicon]AmpereOne Hides Memory Tags in the Bits a Server Already Pays For#053[silicon]Li Auto Lets the Compiler Drive Data Instead of Rebuilding a GPU Cache#054[silicon]Microsoft's Rowhammer Defense Sleeps Until One Sub-Bank Looks Dangerous#055[ai-infra]Reasoning Models Hit the Memory-Capacity Wall Before the Compute Ceiling#056[arch]SPEC CPU 2026 Stops Pretending Every Core Runs the Same Program#057[memory]The SSD Stayed Local, but Its Control Path Moved Three Times#058[memory]A Thousand Cartridges, Four Drives: Why Tape Cloud Storage Must Wait on Purpose#059[ai-infra]Meta Gives Kernel Optimization a Search Tree and a Memory#060[memory]The Hierarchy Became the Product: 34 Industry Memory and Storage Papers from 2025#061[fabric]Ask the Switch Which Path the Probe Actually Took#062[silicon]Precision Moves Into the Architecture: Industry Circuits in 2025#063[fabric]A Local Queue Cannot See the Congestion Two Switches Away#064[fabric]The Training Job Already Knows Which Network Paths Matter#065[arch]A CPU Benchmark That Meta Was Willing to Buy Servers With#066[ai-infra]Why Llama 3 Needed Four Kinds of Parallelism at Once#067[ai-infra]Alibaba Makes Collective Communication Diagnose the Cluster#068[ai-infra]The Largest GPU Jobs Fail Most, but Small Jobs Still Set Fleet Policy#069[ai-infra]AI Energy Needs a Meter That Survives Nine Orders of Magnitude#070[silicon]What Industry Put on Silicon in 2024#071[memory]Memory Stopped Being a Component: 35 Industry Papers from 2024#072[silicon]MI300A Turns the Accelerator Package into the Compute Node#073[silicon]FuriosaAI Replaces the Fixed Matrix with a Shape-Shifting Tensor Engine#074[arch]Vera Rebalances the CPU: 88 Cores Fed by 1.2 TB/s#075[ai-infra]The Agent Graph Becomes the Cloud Scheduler: Murakkab Optimizes the Whole Workflow#076[memory]A Cache Hit That Arrives Late Is Still a Miss: Strata Schedules Long-Context Memory#077[ai-infra]Fill the Idle Half of On-Policy RL: Weave Co-Schedules Rollout and Training#078[ai-infra]When Storage Understands the Job: AITURBO Turns AI I/O into a Group Operation#079[ai-infra]The Datacenter Became the Runtime: Nine Industry Papers That Redefined AI Infrastructure#080[ai-infra]Idle Is Not Available: The Accounting Problem Inside Alibaba's 155,410-GPU Fleet#081[ai-infra]A Silent GPU Error Needs Three Different Questions#082[ai-infra]Trace the Request That Missed Its SLO, Not Every Request Around It#083[ai-infra]The Pipeline Is No Longer Made of Identical Bricks#084[ai-infra]Move the Training Job Before It Learns That a Machine Left#085[ai-infra]The Weather File Expires Before the Datacenter: Google Reprices Cooling for 2044#086[ai-infra]Power Is the Cluster: What 83,000 GB200s Teach About 150 MW#087[ai-infra]What a Neocloud Number Proves, and What It Leaves Unmeasured#088[fabric]An AI Fabric Needs Three Views of the Same Failure#089[ai-infra]When Thousands of GPUs Wait for Storage: Reading MLPerf Storage 2.0#090[arch]One Scheduler for Ten Accelerators: XSched's Three Levels of Preemption#091[ai-infra]Three Percent Globally, a Grid Problem Locally: Reading the IEA's AI Forecast#092[memory]Mobile Memory Takes the Server Socket: LPDDR6 Silicon, SOCAMM2 and DRAM That Computes#093[memory]NAND Moves Up a Tier: Reading the First High Bandwidth Flash Specification#094[silicon]The PCIe Tax on CXL, and the 4 nm Silicon Built to Repeal It#095[silicon]Dropping the PHY: d-Matrix Bonds Compute Directly onto DRAM#096[fabric]Who Wires the Rack: Reading UALink Against NVLink and Ethernet#097[ai-infra]Seven Models, One GPU: What the Long Tail Costs and How Aegaeon Reprices It#098[ai-infra]Evict First, Ask Questions Later: How ByteDance Keeps 9,600 GPUs Worth Training On#099[fabric]The FLOPS You Cannot Buy, You Wire: Inside Tencent's Half-Million-GPU Fabric#100[fabric]Firing the Switch: A Scale-Up Domain Built from Transceivers#101[ai-infra]The Hardware Wishlist DeepSeek Wrote on a Halved NVLink#102[silicon]No HBM, No Apologies: MTIA 2i and Meta's Productionization Ledger#103[ai-infra]One Controller Above, Many Below: HybridFlow, the Engine Inside verl#104[ai-infra]Buying Back FLOPs with DRAM: Mooncake, the Cache That Serves Kimi#105[memory]CXL Leaves the Lab: How Meta Turned Retired DDR4 into a Production Memory Tier#106[memory]KVCache Without the Network: Alibaba Runs LLM Serving on a Real CXL 2.0 Switch