这个基准测试戳中了AI科研助手的关键短板——无法判断研究想法的可行性,做自动化科研或依赖LLM审稿的团队值得关注,看完会重新评估AI在科研流程中的角色。
研究人员推出了SoundnessBench基准测试,包含1,099个从ICLR投稿中重建的机器学习研究提案,并附有评审员的合理性评分。测试了12个前沿大语言模型后发现,它们普遍存在乐观偏差,在标准提示下常将低合理性提案评为合理。即使采用激进提示,也只是将错误从假阳性转为假阴性。控制实验排除了公共语料污染、表面特征等单一干扰因素。结果表明,当前LLM尚不能可靠地作为科学严谨性的独立初审评估者。
SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?
Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. However, existing benchmarks rarely test a fundamental bottleneck: whether Large Language Models can judge the methodological viability of a research idea before expending time and computational resources. We introduce SoundnessBench, a curated benchmark of 1,099 machine-learning research proposals reconstructed from ICLR submissions, labeled with reviewer soundness sub-scores, and audited against source papers. SoundnessBench should be interpreted as a benchmark for recoverable proposal-stage soundness rather than exact prediction of full-paper review outcomes. Across 12 frontier LLMs, we find a pervasive optimism bias: under standard prompting, models frequently rate low-soundness proposals as sound, while aggressive prompting largely shifts errors from false positives to false negatives. Additional controls for public-corpus contamination, paper-identifying phrases, surface features, and human audit quality suggest that this behavior is not explained by a single confounder. Our results indicate that current LLMs are not yet reliable as standalone first-gate evaluators for scientific rigor.