精选理由
NVIDIA 刚开源的 75B MoE 模型,实际只激活 9.3B 参数,单卡就能跑百万 token 上下文,适合长文档或代码分析。
NVIDIA 开源了 NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4 模型,总参数量 75.3B,激活参数 9.3B,采用 MoE 架构。该模型从 Nemotron-3-Super-120B 通过 Iterative Puzzle 框架压缩而来,支持 1M token 上下文长度。模型可运行在单张 GB10 显卡上。
原文 · NVIDIA AI
You're welcome 👊
You're welcome 👊 mr-r0b0t @mr_r0b0t @NVIDIAAI just gifted us a 75B MoE 🤩🤩🤩 nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4 75.3B total / 9.3B active compressed from Nemotron-3-Super-120B using the Iterative Puzzle framework. 1M token context support! Perfect for your single GB10 ♥️ 🔗 View Quoted Tweet 💬 5 🔄 3 ❤️ 88 👀 6694 📊 12 ⚡