OpenAI刚发布的GPT-6 Astra在科学基准测试中得分暴涨42%,远超前代模型。
GPT-6 Astra在Terminal-Bench-Science 0.1基准测试中得分为64.6%。这一成绩比GPT-5.6的22.4%提升了42.2个百分点。Terminal-Bench-Science基准测试发布仅一周就被OpenAI采用。该基准测试旨在解决世界 toughest的科学挑战。
The most capable models available in the world are using Terminal-Bench-Science to show it (just one...
The most capable models available in the world are using Terminal-Bench-Science to show it (just one week after the benchmark’s release). Congratulations to @StevenDillmann and team for their remarkable impact! Steven Dillmann @StevenDillmann What a week for Terminal-Bench-Science & AI for Science in general! 🥂 Just 1 week after launch, Terminal-Bench-Science 0.1 is now also the #1 featured benchmark on today's GPT-6 Astra release by @OpenAI . GPT-5.6 Sol → 22.4% GPT-6 Astra → 64.6% The +42.2pp jump is out of this world, and that's on top of the Claude Fable 5.1 numbers from two days ago. To all scientists out there: We need to buckle up and bring the world's toughest scientific challenges together for Terminal-Bench-Science 0.2. Deadline is October 5. Let the games begin! openai.com/index/gpt-6-as… U 🔗 View Quoted Tweet 💬 1 🔄 3 ❤️ 15 👀 3609 📊 3 ⚡