精选理由
Ben Davis 在 x.com 上展示了 DeepSWE 在多项任务中的出色表现,比 Fable 和 GPT-5.6-sol 更胜一筹,值得关注。
Ben Davis 在 x.com 上分享,DeepSWE 在 10 项任务中表现优于 Fable 和 GPT-5.6-sol,得分超过 80%,引发热议。
原文 · orange.ai
牛来牛逼了,deepswe 大幅超过 fable??? https://t.co/DpHaKw9mYW
牛来牛逼了,deepswe 大幅超过 fable??? x.com/davis7/status/… Ben Davis @davis7 I ran this thing through 10 tasks on DeepSWE (so there could be a ton of variance in it's real score, this is a subset), but uh... gpt-5.6-sol: 52% fable: 65% whatever the hell this is: 80% (was a near miss on the "x"s so actually over 80%) I am very confused 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 0 👀 141 ⚡