做科研的 AI 用户终于有了专门评估 AI 辅助科研能力的基准——T-Bench Science 直接面向真实工作流,科学家可以贡献自己的流程来推动模型进步,值得关注和参与。
Terminal-Bench 是一个评估 AI 模型在计算机上使用工具(如命令行)达成目标能力的基准。现在它扩展到了科学领域,推出 T-Bench Science,专门评估 AI 在真实科研工作流中的表现。该基准面向生命科学、物理、地球科学、数学等领域的科学家,并开放任务贡献至 2026 年 8 月。贡献的科研工作流越多样,越能推动下一代 AI 模型更好地辅助日常研究工作。这不是训练数据集,而是用于评估前沿模型性能的基准。Anthropic、OpenAI 和 Google DeepMind 已使用 Terminal-Bench 评估 AI 编程能力,现在科学领域也加入其中。
I'm very excited about this extension to the celebrated Terminal-Bench to science. If you're a scie...
I'm very excited about this extension to the celebrated Terminal-Bench to science. If you're a scientist (life, physical, earth, mathematical science, etc) interested in AI, definitely check this out! Terminal bench evaluate how good AI models are at controling tools on a computer to achieve a goal (using the command line). T-Bench science now extends that to "AI for Science" and it comes with a call to contribute your own (real scientific world) workflow to the benchmark (until August 2026). The more workflows and the more diverse they are, the better the next generation of AI models will be at helping you in your daily research work. Note that this is not a training dataset, it's to evaluate frontier model performances. Steven Dillmann @StevenDillmann 📣 Announcing Terminal-Bench Science: benchmarking AI agents on real scientific workflows – now open for task contributions👇 tbench.ai/news/tb-scienc… Vt @AnthropicAI , @OpenAI , and @GoogleDeepMind use Terminal-Bench to evaluate AI on coding tasks. We're now extending it to scientific workflows. 1/6🧵 🔗 View Quoted Tweet 💬 2 🔄 1 ❤️ 4 👀 236 📊 2 ⚡