开放模型在评测中访问外部世界
AI21Labs实验发现开放模型通过访问外部世界能大幅提升基准测试成绩,引发对AI评测标准的思考。
AI21Labs让开放模型在公共基准任务评测中访问外部世界。大多数模型找到了包含任务修复程序的上游提交。所有模型在找到修复后得分更高,GPT 5.6 Luna从0.40提升至0.72,GLM-5.3从0.60提升至0.84,MiniMax M3从0.31提升至0.63。评测结果引发对评测内容的质疑:是在测试编程能力还是寻找答案的能力?评测基于SWE-rebench、SWE-smith和Atlas QnA等基准任务。
We let open models access the outside world during evals on tasks from public benchmarks. Here's what happened:
Most went and found upstream commits containing fixes for the tasks they were supposed to solve.
And across every model, scores were higher when the agent found the fix. For some, that gap is huge:
• GPT 5.6 Luna: 0.40 → 0.72 • GLM-5.3: 0.60 → 0.84 • MiniMax M3: 0.31 → 0.63
Introduces big questions about what evals are measuring: ability to code, or ability to find an answer?
Evaluated on tasks from SWE-rebench, SWE-smith, Atlas QnA and more.
- MiniMax_AI09-27 17:06原文