Agon: 竞争性跨模型强化学习的隐式对手评分推理

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

精选理由

Agon让两个模型互相打分比赛,推理能力直接翻倍,比GRPO强一倍,而且完全不需要人工标注过程标签。

AI 摘要

Agon是一种竞争性跨模型强化学习方法,两个模型互为评分器,无需过程标签或奖励模型。在DeepMath硬数据集中使用Qwen3基准,Agon的pass@1达到GRPO的两倍。该方法通过交替角色训练,使模型逐步面对更强的对手。在Qwen3.5和Gemma 4等模型家族的编程代码任务上也验证了有效性。

原文 · arXiv cs.LG

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for good thinking exists. We introduce Agon, which makes two competing models each other's graders. Both attempt the same problem; in alternating roles, one drafts a solution and the other reads it while solving, and each is rewarded for out-solving the other. To win, a model must out-reason a rival that has seen its work, so reasoning is judged implicitly during training, with no process labels and no reward model. Because both models are optimized, each faces a progressively stronger rival, which single-model RL cannot provide. The two need only be comparably strong and behaviorally different. At inference the pair deploys as it trains, a two-stage cascade in which one model drafts and the other answers after reading the draft. On the hard split of DeepMath with Qwen3, this doubles GRPO's pass@1, roughly eight times the gain of an untrained Mixture-of-Agents pass over the same base. The ordering replicates on competitive-programming code and across model families (Qwen3.5, Gemma 4). For now the models talk in text; the next step is to let them reason together in latent space.