论文76°

MIT与哈佛新论文:LLM无法真正进行科学发现

Yet another paper argues that LLMs aren’t close to doing real discovery.

精选理由

MIT和哈佛发了个评测,把前沿LLM丢进真实科研流程,结果它们连假设都提不好。

AI 摘要

MIT和哈佛发布论文《Evaluating Large Language Models in Scientific Discovery》,提出名为SDE的新评估框架。该框架将LLM置于假设提出、实验设计和结果解读的完整研究循环中。测试覆盖生物学、化学、材料科学和物理学的前沿模型。结果显示,LLM在标准科学基准上的表现与真实研究能力之间存在巨大差距。研究者还发现,单纯扩大模型规模无法缩小这一差距。

原文 · Gary Marcus

Yet another paper argues that LLMs aren’t close to doing real discovery.

Yet another paper argues that LLMs aren’t close to doing real discovery. How To Prompt @HowToPrompt__ MIT and Harvard argue LLMs are nowhere near doing real scientific discovery. They published a paper called “Evaluating Large Language Models in Scientific Discovery.” Every week, tech labs claim an LLM has made a breakthrough in biology, physics, or chemistry. But this proves they are faking it. For years, AI benchmarks have tested models using static, multiple-choice science trivia. Models ace these tests, leading everyone to believe AI is right on the verge of autonomous scientific discovery. Researchers built a new evaluation framework called SDE to test what happens when you take LLMs out of the multiple-choice quiz and put them into real, open-ended research projects. They tested frontier models across biology, chemistry, materials science, and physics. The results are sobering. When forced to handle the actual loop of discovery—proposing a testable hypothesis, designing simulations, running experiments, and interpreting ambiguous results iteratively, current LLMs fall apart. There is a massive, glaring performance gap between passing standard science benchmarks and doing real science. Why do they fail? Because real science requires iterative reasoning, handling imperfect evidence, and adapting to unexpected observations. LLMs are built to predict the next token based on existing internet data. They can regurgitate a textbook explanation of photosynthesis or quantum mechanics instantly. But when placed inside an uncharted loop where the textbook doesn't have the answer yet, they hit a wall. Worse still, the researchers discovered diminishing returns. Simply scaling up model sizes and adding raw compute isn't fixing the gap. Top-tier models from different providers share the exact same blind spots. We are miles away from general scientific superintelligence. The tech industry is selling a narrative that AI is about to automate labs, run clinical trials, and invent materials on autopilot. But right now, AI isn't doing science. It's just remembering it. 🔗 View Quoted Tweet 💬 2 🔄 1 ❤️ 1 👀 1001 📊 2 ⚡

MIT与哈佛新论文:LLM无法真正进行科学发现 · AI 热点