预测者普遍低估 AI 进展:IMO 金牌比专家预测提前五年
把专家和超级预测者的 AI 预测跟实际成绩逐条对账,IMO 金牌提前了五年,数字都在这,看看大家都错在哪。
这份盘点对比了专家与超级预测者(SF)对 AI 能力的预测与实际结果。OpenAI 在 2025 年 7 月达到 IMO 金牌水平,比专家中位预测的 2030 年提前约五年;MATH 87.8% 和 MMLU 88.7% 的成绩也远超 2022 年时专家 21%/25% 的概率估计。Cybench 90% 于 2026 年达成,早于专家预测的 2028 年。偏差最大的是 LiveCodeBench Pro Hard,2026 年 5 月实际达到 53.8%,而预测仅为 12%-14%。生物风险则相反,实际仅 5%(RCT),低于超级预测者的 16% 和病毒学家的 40%。
Forecasters were mostly wrong about AI.
On capabilities
- Claimed Millennium Prize problem solution (September 2026 claim by OpenAI re Navier–Stokes, targeting end of 2027, awaiting verification from prize organization; forecast: 10% [experts] and 5% [SF]) - IMO gold (July 2025, forecast: 2030 [experts] and 2035 [SF]), virology troubleshooting task comparable to a strong team of humans (April 2025, forecast: 2030 [experts] and 2034 [SF]) - 90% on Cybench (2026; forecast: 2028 [experts] and 2030 [SF]) - MATH score of 87.8% and MMLU score of 88.7% (mid-2024, forecast: 21%/25% probability [experts] and 9%/7% [SF] as of 2022) - QuALITY hard-subset score of 69.3% (2024 target met a year early; forecast: 44% probability [experts] and 20% [SF])
The biggest misses were IMO gold, which occurred five years before experts’ median estimate, biorisk (which resolved much lower, at 5% (RCT), than expected by SFs (16%), biosecurity experts (23%), and virologists (40%)), and Navier-Stokes, which is not yet officially confirmed.
The biggest misses in terms of percentage points have been LiveCodeBench Pro Hard (actual: 53.8% in May 2026; forecast: 12%-14% by the end of 2026) and FrontierMath Tiers 1-3 unadjusted (actual: 40.7% as of late 2025; forecast: around 30%-31%). Superforecasters were more skeptical than exper ts on the timing of technical achievements.
SF = Superforecasters