论文73°

腾讯混元联合 LMArena 推出 WebCraftBench 基准

Real-world app requests don’t come as perfect specs. Congrats to the @TencentHunyuan team on WebCraf...

精选理由

腾讯混元做了个新基准 WebCraftBench,用 369 条真实需求测 AI 写网页到底靠不靠谱,测了 17 个模型,和 Code Arena 榜单相关性 0.89。

腾讯混元团队基于 Code Arena 内部复刻环境收集的 369 条真实请求,构建了网页生成基准 WebCraftBench。该基准覆盖口语化描述、指令不完整、需求范围不一等真实使用场景,让智能体实际操作生成的应用并进行评分。基准对 17 个模型排名,与 Code Arena 排行榜的相关性达 0.89。在 197 组人工验证样本上,其评分与人类偏好一致率为 85.3%。

图片来源 · lmarena.ai
原文 · lmarena.ai

Real-world app requests don’t come as perfect specs. Congrats to the @TencentHunyuan team on WebCraf...

Real-world app requests don’t come as perfect specs. Congrats to the @TencentHunyuan team on WebCraftBench! Built from 369 requests collected from their internal replica of Code Arena, it embraces the messy details of real human usage: informal language, incomplete instructions, different project scopes, and diverse preferences. Their benchmark tests how AI-built apps look, work, and deliver on those requests. Across 17 models, its rankings strongly correlate with our Code Arena leaderboard (0.89). What could you uncover from real human-AI conversations and Arena’s leaderboard history? Explore our open datasets at huggingface.co/lmarena-ai/dat… ! Tencent Hy @TencentHunyuan The AI says the site is done. Then the homepage errors, the buttons overlap, and it added a login flow you never asked for. 🙃 Introduces WebCraftBench: agents actually use the live app, coverage-guided exploration finds what never got reached, then we score aesthetics, usability, and whether the original request was met. On 197 human-validated pairs, it matches human preference 85.3% of the time. Paper: arxiv.org/abs/2609.15387 Z 🔗 View Quoted Tweet 💬 0 🔄 0 ❤️ 1 👀 1237 ⚡