揭露闭源API的基准测试猫腻
Hugging Face CEO Clement Delangue在推特上回应关于模型基准测试的争议,指出闭源API可以通过回退(fallback)机制提升分数,例如Fable 5模型回退到Opus 4.8可能获得更高总分,即使Opus 4.8平均分更低。他以AA基准中的GPQA Diamond和AA-Omniscience为例,说明模型对不同查询的表现不一致,导致回退策略可能掩盖真实能力。Delangue强调,只有API提供商知道实际路由策略,这使得基准测试缺乏透明度。
To people in the answers saying "but opus 4.8 is weaker so without fallback, the score would even be...
To people in the answers saying "but opus 4.8 is weaker so without fallback, the score would even be higher": this is not necessarily true because of how any benchmark - which is an average of queries - work and what is called "the fallacy of division". Even if Opus 4.8 has a lower average score on AA than Fable 5, it actually performs better than Fable 5 on some benchmark that compose the index of AA, especially where there's high refusal rate of Fable 5 (ex GPQA Diamond, AA-Omniscience). The same would go if you'd take a single benchmark btw as it's always an average of queries and the fact that a model has a higher score on average doesn't mean they answer better on 100% of queries. So it's possible that Fable with Opus 4.8 fallbacks is getting a higher score than pure Fable, even if Opus 4.8 is weaker on average. The challenge is no one knows, except the API provider, which is the challenge I'm pointing out. More details below from Fable (or Opus?) themselves! clem 🤗 @ClementDelangue This graph captures what’s broken about AI evals: they structurally favor closed-source APIs that can route, fallback, ensemble, and optimize behind the scenes with no transparency. No offense, @ArtificialAnlys , but how is comparing one model to two models fair? 🔗 View Quoted Tweet 💬 6 🔄 3 ❤️ 25 👀 3046 📊 9 ⚡