Nvidia 发布光子共封装光学交换机视频,降低 AI 数据中心功耗

Nvidia released this video of its photonics co-pac…

精选理由

Nvidia 的 CPO 技术直接解决了 AI 数据中心网络功耗和故障率两大痛点,做大规模 GPU 集群部署的团队值得关注,能显著降低运营成本。

AI 摘要

Nvidia 发布了一段关于其光子共封装光学(CPO)交换机的视频,展示了 Lambda 技术。CPO 将光通信组件直接集成到网络芯片附近,取代了传统的可插拔模块,从而大幅降低功耗并减少故障点。在 128,000 GPU 的数据中心中,传统方案需要约 655,000 个可插拔收发器模块,而 CPO 彻底消除了这一组件类别。对于智能体工作负载,CPO 能提供弹性、高效的数据传输,避免 GPU 等待数据,提升推理效率。

原文 · rohanpaul_ai

Nvidia released this video of its photonics co-pac…

Nvidia released this video of its photonics co-packaged optics (CPO) switch with Lambda.

The AI race is not only about stronger GPUs, but about wasting far less power while those GPUs talk to each other.

With co-packaged optics (CPO), NVIDIA is putting the light-based communication parts much closer to the main networking chip, instead of placing them as separate plug-in modules at the edge of the switch.

From NVIDIA's official blog on this

"co-packaged optics (CPO) connects directly to the token economy. Network power is overhead: it keeps GPUs connected but doesn't generate tokens. Network failures are also overhead: they turn provisioned GPU capacity into idle capacity. CPO addresses both by reducing network power draw and removing a large class of pluggable optical components from the fabric.

A 128,000-GPU data center using traditional pluggable transceivers requires roughly 655,000 discrete transceiver modules across the switching fabric. Each one is a potential failure point. CPO removes that component class entirely.

Agentic workloads change the pressure on the network. A traditional inference request is relatively self-contained. An agentic request can involve planning, retrieval, tool use, multiple model calls, and follow-up reasoning. More data moving across the cluster. More points where network latency or failure affects the outcome.

Multi-agentic inference needs elastic and resilient data movement, so GPUs are not waiting for data, while maintaining tokens per second and fast time to first token."