技巧精选73°

同一模型在不同代理框架中得分差异显著

精选理由

HuggingFace开源了多框架强化学习方法,让同一模型在不同代理框架中表现提升12%,还减少了31%的工具调用次数。

HuggingFace团队发布了多框架强化学习指南。同一模型在一种代理框架中得分为62%,在另一种中仅为33%。研究人员通过代理方法训练LFM2.5-2.6B模型,使其在四个框架中的平均得分从42%提升至54%。OpenCode框架单独训练使模型得分从34%提升至58%,而多框架模型在所有框架上都有改进。

原文 · Hugging Face

The same model, with the same weights, scores 62% in one agent harness and 33% in another. @adithya_s_k and the @huggingface team just released the ultimate guide to multi-harness RL, and it's one of the most practical RL write-ups this year, and everything open! The trick is simple. Don't touch the harness. Point it at a proxy instead of the model. The proxy speaks all four API formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini). It records the exact token ids and logprobs vLLM sampled, and you train on that. You don't change a single line of Claude Code, Codex or OpenCode. Results: 🔹 Trained across 4 harnesses at once, LFM2.5-2.6B by @liquidai went from 42% to 54% 🔹 31% fewer tool calls, thanks to a small bonus for solving tasks in fewer steps 🔹 Training in OpenCode alone took OpenCode from 34% to 58%, but the multi-harness model improved everywhere They also tried the shortcut everyone reaches for: fine-tune on 3,189 successful rollouts from Qwen3.8-27B. Imitation plateaued at 47.5%, below both RL runs. Copying a bigger model doesn't get you there. Practice does. The best part is that everything is open: the capture proxy in OpenEnv, the trainer in TRL, the tasks, the SFT data, the training code and all seven trained models. Agents will run in dozens of harnesses. Now open models can be trained for each of them, by anyone. Read it here huggingface.co/spaces/FineEnv… 9SLS 💬 12 🔄 5 ❤️ 29 👀 3587 📊 17 ⚡