NVIDIA把这个MoE模型压到75B,吞吐翻倍,用户延迟不变,适合降本增效。
NVIDIA发布Nemotron-Labs-3-Puzzle-75B-A9B,这是Nemotron-3-Super的压缩变体。通过Iterative Puzzle方法,总参数从120.7B降至75.3B,活跃参数从12.8B降至9.3B。在单台8xB200节点上,它实现了2.03倍于Super的总吞吐量,每用户100 tok/s。在H100上,1M-token并发从1提升到8。
NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput
NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. Iterative Puzzle alternates hardware-aware structural compression with short knowledge distillation recovery phases. The model drops from 120.7B total / 12.8B active parameters to 75.3B / 9.3B. On a single 8xB200 node it delivers 2.03x Super's total throughput at 100 tok/s per user. On one H100, 1M-token concurrency rises from 1 request to 8. The post NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput appeared first on MarkTechPost .