AI模型精选

PrimeIntellect发布 prime-rl 0.6.0:基于vLLM的万亿级Agentic RL训练

What does trillion-scale agentic RL look like on t…

精选理由

PrimeIntellect 用 vLLM 把 agentic RL 搞到万亿参数级了,FP8 加专家并行,28 个 H200 节点跑 131k 序列,每步不到 5 分钟,训练 GLM-5 做 SWE 任务。

AI 摘要

PrimeIntellect 发布了 prime-rl 0.6.0,在 vLLM 上运行万亿级 agentic RL 推理。该版本采用 FP8 量化、专家并行、prefill/decode 分离以及 KV 缓存卸载(原生+Mooncake),并使用 vllm-router。在 28 个 H200 节点上,以 131k 序列长度训练 GLM-5 执行 SWE 任务,每步耗时不到 5 分钟。@m_sirovatka 将在 vLLM Office Hours 中深入讲解。

图片来源 · vLLM
原文 · vLLM

What does trillion-scale agentic RL look like on t…

What does trillion-scale agentic RL look like on the inference side?

@PrimeIntellect's prime-rl 0.6.0 runs it on vLLM — FP8, wide expert parallelism, prefill/decode disaggregation, KV cache offloading (native + Mooncake), and vllm-router — to train GLM-5 on SWE tasks at 131k seqlen with sub-5-min steps on 28 H200 nodes.

📅 @m_sirovatka goes deep live in vLLM Office Hours today.

📖 Read up ahead of time: https://t.co/pW50ai35C4