模型

Allen AI 的 SciArena 基准测试结果:1700 用户近 4000 投票

精选理由

Allen AI 用研究者投票测了 AI 的科学文献能力,1700 人投了 3900 票,想看看哪个模型更靠谱?这里直接给结果。

Allen AI 构建的 SciArena 基准测试用于评估 AI 模型处理科学文献问题的能力,评判者均为研究人员。该基准将于 7 月 15 日退役,最终数据包括约 1700 名用户投出的约 3900 次投票。结果揭示了模型在科学问答、文献理解等任务上的表现对比。具体排名和用户偏好信息已在推文线程中公布。

原文 · Allen AI (Ai2)

We built SciArena to test how well AI models handle scientific literature questions, as judged by researchers.

It's retiring July 15, and the results are in: ~1,700 users cast ~3,900 votes.

Here's what they told us. 🧵 https://t.co/sU1JYfXOhU