GameDevBench基准测试发布
Game development remains one of the most-requested, and most-challenging, categories on Arena. @iam...
CMU博士生发布游戏开发基准,测试AI在1小时内能完成多少人类新手任务
Carnegie Mellon大学博士生Wayne Chi在Arena平台发布了GameDevBench基准测试。该基准基于真实教程构建,将游戏开发转化为可验证的确定性任务。测试评估了前沿模型在人类初学者1小时内可完成的任务上的表现,分析了编程与多模态理解哪个是最大瓶颈。
Game development remains one of the most-requested, and most-challenging, categories on Arena. @iam...
Game development remains one of the most-requested, and most-challenging, categories on Arena. @iamwaynechi , PhD candidate at Carnegie Mellon University and research intern at Arena, just walked us through GameDevBench: a benchmark built from real tutorials that turns game development into verifiable, deterministic tasks. How do top frontier models perform on tasks a human beginner could complete in under an hour, and is the biggest bottleneck coding or multimodal understanding? Check out Wayne’s full talk to hear results and learn more about GameDevBench: youtube.com/watch?v=yoHcvr… Your browser does not support the video tag. 🔗 View on Twitter 💬 2 🔄 3 ❤️ 27 👀 3521 📊 6 ⚡