NVIDIA出了个轻量模型,30B参数但只激活3B,跑长任务上下文能到1M,单卡就能部署,速度还快4倍。
NVIDIA推出Nemotron 3.5 Lightning,专为长时运行的智能体任务设计。该模型总参数量30B,激活参数3B,支持高达1M token的上下文,并已开放商用。NVIDIA提供BF16和更小的NVFP4量化检查点,支持在单张H100或DGX Spark上部署,也兼容RTX 5090。通过多token预测、DSpark和DFlash投机解码以及NVFP4量化,生成速度最高提升4倍。在PinchBench基准上,其准确率达86%,完成速度比Qwen3.6 35B快30%。
源:https://t.co/TCFUzoABCI
源: x.com/rohanpaul_ai/s… Rohan Paul @rohanpaul_ai NVIDIA released Nemotron 3.5 Lightning for long-running agent execution. - 30B total / 3B active parameters, with up to 1M tokens of context. Ready for commercial use. - Local deployment: NVIDIA ships BF16 and much smaller NVFP4 checkpoints, lists 1× DGX Spark or 1× H100 for single-GPU deployment, and also lists RTX 5090 among supported hardware. - NVIDIA then adds multi-token prediction, DSpark and DFlash speculative decoding, plus NVFP4 quantization to push generation speed further. - NVIDIA claims up to 4× output speed; on PinchBench, 86% accuracy and 30% faster completion than Qwen3.6 35B at similar accuracy. 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 256 ⚡