Meta研究:双代码审查代理减少静默错误
Meta论文揭示双代理审查比单代理预算提升效果更好,Claude Code和Codex组合错误率降低。
Meta新论文发现,使用两个不同的代码代理互相审查补丁比单个代理获得更大预算能捕获更多静默错误。混合使用Claude Code和Codex将完全正确补丁比例从45.8%提升至62.5%。代理编辑真实训练代码可能导致测试数据泄露、梯度断裂或标志错误。
New Meta paper finds that having 2 different coding agents review each other's patches catches far more silent bugs than giving 1 agent a bigger budget.
Mixing Claude Code and Codex on the same task raised fully correct patches from 45.8% to 62.5% at matched spend.
Agents editing real training code can leak test data, break a gradient, or miswire a flag.
The code still runs, so you burn GPU hours and get numbers that look valid.
The reason is uncorrelated mistakes: agents from the same product fail the same way, so there is little for review to catch. The effect repeated on an unrelated training codebase.
– arxiv. org/abs/2609.39551
Title: "RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models"