AgileRL 把 Nemotron 3.5 Lightning 放到 Arena 上微调,6 小时就能超过 Claude Sonnet 5,训练完权重还归你。
AgileRL 宣布与 NVIDIA 合作,支持 Nemotron 系列模型的后训练。NVIDIA 提前提供了 30B3A 参数 MoE 模型 Nemotron 3.5 Lightning。在 Arena 平台上,该模型经微调后在长程推理和真实客服任务上超过 Claude Sonnet 5。单次训练只需四块 H100 GPU 运行六小时即可掌握推理任务,上下文长度超过 50,000 token。Arena 客户可在 Nemotron 模型发布后立即获得相应的训练与评估环境。
https://t.co/EqHNOzY2q7
x.com/AgileRL_Inc/st… AgileRL @AgileRL_Inc We are proud to announce a collaboration between AgileRL and @NVIDIAAI to support the post-training of Nemotron models. NVIDIA's Nemotron open-source model family is now available for fine-tuning on Arena, our platform for creating AI agents specialized at any task. NVIDIA gave us early access to their latest 30B3A parameter MoE model, Nemotron 3.5 Lightning. On Arena, we trained it to outperform Claude Sonnet 5 on two different tasks: a long-horizon reasoning challenge and a real customer support workload. Nemotron proved exceptionally easy to post-train: a single six-hour run on four H100 GPUs was enough to master the reasoning task, at context lengths beyond 50,000 tokens. The full results are in the announcement, linked below. With this collaboration, Arena customers get access to NVIDIA Nemotron models as soon as they are released, with the training and evaluation setup already built around them. Post-train Nemotron on your own data, environment and edge cases, to build an agent that masters your task. Across industries including banking, insurance, aerospace and government, production traffic is moving off frontier model APIs and onto infrastructure these companies control. Businesses want agents with genuine expertise in their specific task, trained on their own data. Security, control and data sovereignty demand that model weights remain on their own infrastructure. And they want the fixed cost of hardware they own, rather than per-token pricing that compounds with every request. This collaboration provides that path. Begin with a dataset, an RL environment, or simply a description of the job the agent has to do, and our team will build the rest with you. Arena handles the training and deploys the finished agent in one click, with the weights yours to keep. 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 1 👀 315 ⚡