论文精选

AI时代数据中心电力层级设计:应对1MW机架密度挑战

Designing Datacenter Power Delivery Hierarchies for the AI Era

精选理由

数据中心电力设计是AI基础设施的瓶颈,这篇论文用微软Azure的实际数据量化了电力搁浅的代价,做数据中心规划或AI硬件部署的团队值得一读。

AI 摘要

随着AI加速器需求激增,数据中心机架功率密度预计到2027年将接近每部署1MW,这对电力输送设计构成重大挑战。传统数据中心若针对不同密度目标设计,可能导致电力搁浅,即无法充分利用已配置的电力容量。论文提出一个评估框架,结合GPU、计算和存储部署的投影模型与微软Azure的生产数据,分析多资源搁浅对可部署容量、资本支出和性能的影响。结果表明,规划目标不应是装机兆瓦数,而是随时间变化的可部署容量。该框架帮助设计者在长期运营中保持效率,适应多代硬件和不断变化的工作负载。

原文 · arXiv cs.AI

Designing Datacenter Power Delivery Hierarchies for the AI Era

Demand for AI accelerators is rapidly increasing rack power density, with projections approaching 1MW per deployment by 2027. This poses a major challenge for datacenter power delivery designers. As power densities increase, a datacenter designed for a different target density may strand power, i.e., may be unable to use all the power that its delivery hierarchy has provisioned. Designs must remain efficient over long datacenter lifetimes and multiple hardware generations. Power utilization is particularly important as grid power capacity is a scarce resource in the AI era. Designing an efficient power delivery hierarchy for the long run is difficult because rack placement feasibility, workload impact, and cost depend jointly on electrical topology, deployment granularity, placement policy, power oversubscription, and workload mix. Moreover, each of these factors evolve over time, have inter-dependencies across multiple resource dimensions, and generally do not lend themselves to closed-form analysis. To address this challenge, we develop a framework for evaluating datacenter power delivery designs using throughput, power, and cost metrics over realistic arrival, oversubscription, and decommissioning sequences. The framework combines projection models for GPU, compute, and storage deployments with operational factors grounded in production data from Microsoft Azure. Our results show that multi-resource stranding materially changes deployable capacity, effective capital expenditure, and delivered performance, and quantify how rising density from rack- and pod-scale AI systems shapes these outcomes. For AI datacenter design, the relevant planning objective is not installed megawatts, but deployable capacity over time.