Hugging Face 把 Claude Code、Codex 等编程 harness 变成 RL 训练环境
Hugging Face 把 Claude Code、Codex 这些真实编程工具直接当 RL 训练环境用,不用改一行代码,LFM2.5-2.6B 训练后 OpenCode 成绩从 34% 涨到 58%,全套开源可复现。
Hugging Face 团队用捕获代理的方式,把 Claude Code、Codex、Hermes、Pi、OpenCode 等 10 个编程 harness 直接变成 RL 训练环境,harness 和训练代码都无需改动。代理支持 OpenAI Chat Completions、OpenAI Responses、Anthropic Messages、Gemini 四种格式,转发给 vLLM 并记录 token ID 和 logprobs 供 TRL 训练。在 Liquid AI 的 LFM2.5-2.6B 上测试:单 harness 训练使 OpenCode 成绩从 34% 升到 58%,4 个 harness 同时训练则全部提升(42% 到 54%),而用 Qwen3.8-27B 的 3,189 条 rollout 做 SFT 只到 47.5%。同一模型在 Mini-SWE-Agent 下得 62%,在 Claude Code 下只有 33%,说明训练环境与实际部署环境一致的重要性。另外给少用工具调用加奖励后,模型在已解决任务上的调用次数减少 31%,Codex 下约减半。
We turned Claude Code, Codex, Hermes, Pi, @opencode and other coding harnesses into RL environments. No changes to the harnesses, no changes to the training code. Any open model, any task set, fully open source my friends! Same model, same weights: 62% under Mini-SWE-Agent, 33% under Claude Code. But training inside a real harness normally means reimplementing it as an environment, so most models get trained in a scaffold nobody actually ships. The fix is a proxy, not a rewrite. The harness thinks it's talking to a model API. It's actually talking to a capture proxy that speaks the 4 formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini), forwards to @vllm_project , records the exact token IDs and logprobs vLLM sampled, and hands TRL sequences it can train on. The harness becomes the environment. 10 harnesses run through it today, none modified. And because you control the reward, you can shape behavior the harness never asked for. We added a small bonus for solving a task in fewer tool calls: on tasks it already solved, the model now uses 31% fewer calls, in every harness, and about half under Codex. Tested on LFM2.5-2.6B from @liquidai : → Train in one harness: better mostly in that harness (OpenCode 34% → 58%). → Train in 4 at once: better in all 4 (42% → 54%). → SFT on 3,189 rollouts from Qwen3.8-27B instead: plateaus at 47.5%, below both RL runs. Everything is open and reproducible: the capture proxy in OpenEnv, the trainer in TRL, the tasks, the SFT data, the training code and all 7 trained models. Bigger models and bigger runs next. Full guide: huggingface.co/spaces/FineEnv… 💬 14 🔄 4 ❤️ 42 👀 3160 📊 19 ⚡