Meta 推出 WildArtifactBench 评估框架,用胜率和 Elo 分数评估智能体,覆盖多模态工作流,还开放了 10 项任务。
Meta 预览 WildArtifactBench,一个内部评估框架。该框架通过人类和智能体偏好评委的胜率和 Elo 分数,而非严格的真实基准,评估智能体在复杂现实任务中的表现。它扩展了任务覆盖范围,涵盖实用的多模态工作流。Meta 将发布 WildArtifactBench 中的 10 项任务。
Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess a...
Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats. By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth rubrics, it expands task coverage across practical multimodal workflows. We’re releasing 10 tasks from WildArtifactBench as a step forward in our ability to measure the real practical utility delivered by multimodal agents: research.meta.ai/wild-artifact-… 💬 9 🔄 7 ❤️ 80 👀 7131 📊 20 ⚡