产品多源确认精选72°

NVIDIA Dynamo Snapshot:Kubernetes 推理工作负载冷启动从分钟级降至5秒

精选理由

Kubernetes 上跑推理的团队终于不用忍受 GPU 空转几分钟了——Dynamo Snapshot 把冷启动压到 5 秒,做弹性扩缩容的 MLOps 工程师可以直接拿来用。

NVIDIA 推出 Dynamo Snapshot,一种针对 Kubernetes 上推理工作负载的快速启动方案。该方案利用 GPU 内存快照(GMS)实现高速互连上的并发权重恢复,同时结合 Linux 原生 AIO 和并行 memfd 恢复技术,加速 CRIU 恢复性能。在推理部署中,需求波动导致冷启动耗时数分钟,造成 GPU 闲置。Dynamo Snapshot 将启动时间从分钟级缩短至 5 秒以内,显著提升 GPU 利用率和推理效率。

原文 · NVIDIA AI

Introducing Dynamo Snapshot, our approach for fast startup for inference workloads on Kubernetes, which reduces startup time from minutes to under 5 seconds. In production inference deployments demand fluctuates over time. Cold-starting inference workloads can take minutes, leaving idle GPUs that generate no tokens and serve no requests. Snapshot leverages GMS to enable concurrent weight restoration over a high-speed interconnect, while using Linux native AIO and parallel memfd restoration to accelerate CRIU restore performance. 💬 6 🔄 8 ❤️ 65 👀 4047 📊 16 ⚡