精选理由
Claude在两项重要基准测试中大幅领先Fable 5,科学任务得分翻倍,综合能力提升13.8%。
Claude模型在Terminal-Bench-Science 0.1测试中得分为52.6%,是Fable 5的两倍多。在Terminal-Bench 4.0测试中,Claude得分为55.8%,而Fable 5为42.0%。这些成绩表明Claude在科学任务和综合能力方面表现优异。
原文 · Claude
Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1...
Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. 💬 100 🔄 331 ❤️ 4046 👀 614643 📊 568 ⚡