Claude模型在多项基准测试中创纪录

Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1...

精选理由

Claude在两项重要基准测试中大幅领先Fable 5,科学任务得分翻倍,综合能力提升13.8%。

AI 摘要

Claude模型在Terminal-Bench-Science 0.1测试中得分为52.6%,是Fable 5的两倍多。在Terminal-Bench 4.0测试中,Claude得分为55.8%,而Fable 5为42.0%。这些成绩表明Claude在科学任务和综合能力方面表现优异。

原文 · Claude

Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1...

Across our benchmarks, the model sets a new standard. It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. 💬 100 🔄 331 ❤️ 4046 👀 614643 📊 568 ⚡

Claude模型在多项基准测试中创纪录 · AI 热点