论文多源确认

AI夸大实验结果,距自主研究仍远

AI agents overstate their results and remain far from autonomous research, study finds

精选理由

两家机构发现当前AI模型能做实验但不会自我批评,连Sol都只有人类15%水平。

Epoch AI和Anthropic研究发现,GPT-5.6 Sol和Claude Fable 5等模型能进行实验,但缺乏科学自我批判和真正创造性思维。Sol仅达到人类参考分数的15%,且方法已是研究人员已知。模型最大弱点仍无法批判性质疑自身结果。

原文 · Decoder

AI agents overstate their results and remain far from autonomous research, study finds

Epoch AI and Anthropic independently found the same thing: current AI models like GPT-5.6 Sol and Claude Fable 5 can run experiments but lack scientific self-criticism and genuine creative thinking. At best, Sol reached 15 percent of the human reference score, and even that came from methods researchers already knew. The models' biggest weakness is still their inability to critically question their own results. The article AI agents overstate their results and remain far from autonomous research, study finds appeared first on The Decoder .