视频生成模型 Video DeltaNet 提升直播视频生成效率
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
这个新方法能显著提升视频生成速度,对需要快速生成直播视频的场景很有用。
视频生成模型 Video DeltaNet 通过结合局部 Softmax 注意力和双向线性内存,解决了视频扩散模型在去噪过程中处理长时空序列时的计算瓶颈问题。其线性分支引入的 Video Delta Attention (VDA) 每帧更新一次记忆,同时整合空间令牌。在 MiniMax H3 模型上应用后,VDN-H3 在八块 NVIDIA B200 GPU 上完成 14.3 秒 768p 视频的去噪仅需 6.70 秒,比 50 步密集 H3 基准快 14.5 倍。
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation
Video diffusion models repeatedly process long spatiotemporal token sequences during denoising, making attention a major computational bottleneck. Linear attention offers an appealing alternative and has been widely adopted in recent large language models, but directly applying it to video models often fails to preserve the fine-grained interactions required for high-quality generation. We present Video DeltaNet (VDN), which combines local Softmax attention with bidirectional linear memory for long-range video context. Its linear branch introduces Video Delta Attention (VDA), which updates memory once per frame by jointly incorporating its spatial tokens. Separate output projections and learnable gates calibrate the two branches, while a staged teacher-alignment recipe progressively introduces the new pathway into pretrained models. We instantiate VDN on MiniMax H3, applying the hybrid to video-to-video interactions while retaining Softmax for interactions involving text or audio. With eight-step distillation and an optimized SGLang serving stack, VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs, corresponding to a 14.5x speedup over the 50-step dense H3 baseline on the same GPU count.