Artificial Analysis 推出 Terminal-Bench-Science 0.1 智能体科研基准榜单
Stanford 联合 Terminal-Bench 团队搞了个科研智能体榜单,70 个真实科研任务里 GPT-6 Astra 拿 63% 排第一,开源模型才 10%,差距很直观。
Artificial Analysis 上线 Terminal-Bench-Science 0.1 排行榜,该基准由 Stanford 的 Steven Dillmann 与 Terminal-Bench 团队等合作构建,包含 70 个专家整理的真实科研任务。任务覆盖生命科学、物理科学、数学科学、工程科学和地球科学五个领域,智能体在沙盒环境中独立完成并按 pass@1 评分。GPT-6 Astra (max) 以 63% 排名第一,Claude Opus 5.5 (xhigh) 以 62% 紧随其后,是仅有的两个超过 50% 的模型。Claude Opus 5.5 (xhigh) 在数学科学任务上通过率 71%,但生命科学只有 46%。开源权重模型差距明显:GLM-5.3 (max) 得分 10%,DeepSeek V4.1 Flash (max) 得分 9%,落后头部模型超过 50 分。
Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62%
Terminal-Bench-Science 0.1 was announced in August 2026, built by @StevenDillmann and researchers at Stanford with the @terminalbench team and a global community of open source scientific contributors. It has 70 expert-curated tasks based on real research in five domains: life sciences (19), physical sciences (17), mathematical sciences (17), engineering sciences (9), and earth sciences (8).
As with Terminal-Bench, each task drops an agent into a sandbox environment with the data, tools and instructions for a task, and the agent runs from start to finish. Automated tests grade every task pass/fail, and we report the average pass@1 over 3 attempts.
Key takeaways:
➤ Only two models score above 50%: GPT-6 Astra (max) at 63%, and Claude Opus 5.5 at 62% (xhigh) and 59% (max). Recent releases have made large advancements, but even the top model, GPT-6 Astra, has significant headroom
➤ Life sciences is the lowest-scoring domain for most leading models, though domain scores are noisy with 8 to 19 tasks each. Claude Opus 5.5 (xhigh) passes 71% of mathematical sciences tasks but 46% of life sciences tasks
➤ The best open weights models, GLM-5.3 (max) at 10% and DeepSeek V4.1 Flash (max) at 9%, sit more than 50 points below the leaders