Sol-H3 加速 MiniMax-H3 推理:云端最高 30 倍提速、边缘单卡可运行
Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge
MiniMax-H3 生成太慢太吃显存?这篇把云端和边缘两端都解决了,最快 30 倍提速,单张 DGX Spark 就能跑视频生成。
论文针对 MiniMax-H3 的 33B 参数和多步去噪带来的计算开销,提出全栈推理管线 Sol-H3。算法上采用跨分辨率两阶段调度:先用低分辨率步骤确定全局布局,再用高分辨率步骤精修细节,并通过学习的 latent-to-latent 映射省去 VAE 解码-重编码。算子层面用 Recursive Self-Improvement 循环自动搜索核融合与内存布局。最终端到端最高提速 30 倍、内存降低 20%:8xGB200 节点上 5 秒 1344x768 带音频视频比实时快 3.5 倍,单张 DGX Spark 上不到一分钟可全量常驻内存生成。
Sol-H3: Recursive Self-Improvement for MiniMax-H3 Inference Acceleration on Sol-Engine across Cloud and Edge
Video diffusion models are rapidly scaling and exhibiting enhanced generation capabilities. Among these recent advancements, MiniMax-H3 stands out as a highly capable, production-level open-source model. However, its 33-billion parameters and multi-step iterative denoising process introduce substantial computational overhead. Consequently, their practical production is hindered by generation latency in the cloud deployment like NVIDIA-GB200, alongside strict memory limits that pose further challenges at the edge device like DGX-Spark. To address these diverse hardware bottlenecks from cloud to edge device, we present a full-stack inference pipeline that integrates efficient algorithmic design with optimized operator implementations. Algorithmically, we introduce a cross-resolution two-stage generation scheduler that exploits the step-wise nature of diffusion: early low-resolution steps rapidly establish the global layout, while later high-resolution steps focus refinements of local and perceptual details. These stages are connected by a learned latent-to-latent mapping module, completely eliminating the computationally expensive VAE decode-reencode cycle for resolution transferring cross different resolutions. For operator implementation, we deploy a Recursive Self-Improvement (RSI) loop that searches kernel fusions and memory layouts, evaluating latency together with numerical agreement. Together, these optimizations deliver up to 30x end-to-end speedup and 20% lower memory: a 5-second 1344x768 video with audio is generated 3.5x faster than real time on an 8xGB200 node, and in under a minute fully memory-resident on a single DGX Spark.