论文83°

研究反驳Anthropic与OpenAI:自主AI研究尚不可及

Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach

精选理由

普林斯顿和英国AI安全所让Claude和GPT-5写论文,全拒,打脸Anthropic和OpenAI。

AI 摘要

AI智能体使用Claude Opus 4.8和GPT-5.6 Sol,在六天、3000美元API预算和GPU访问权限下独立撰写AI研究论文。这些论文交由未发表的NeurIPS论文原作者评审,结果全部为“Reject”。研究显示,Claude Opus 4.8和GPT-5.6 Sol能完成研究工程流程,但缺乏研究判断、创造性问题解决和放弃失败方法的能力。这一结果直接反驳了Anthropic与OpenAI关于自主AI研究即将实现的声明。

原文 · Decoder

Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach

AI agents using Claude Opus 4.8 and GPT-5.6 Sol were given six days, $3,000 in API credits, and GPU access to independently write AI research papers. The original authors of unpublished NeurIPS papers rated the results as "Reject." According to the study, conducted with Princeton and the UK AI Security Institute, frontier models can handle the full research engineering process but fall short on research judgment, creative problem-solving, and the ability to abandon failed approaches. The article Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach appeared first on The Decoder .