Meta 预览 WildArtifactBench 评估框架
Meta 推出 WildArtifactBench 评估框架,用胜率和 Elo 分数评估智能体,覆盖多模态工作流,还开放了 10 项任务。
Meta 预览 WildArtifactBench,一个内部评估框架。该框架通过人类和智能体偏好评委的胜率和 Elo 分数,而非严格的真实基准,评估智能体在复杂现实任务中的表现。它扩展了任务覆盖范围,涵盖实用的多模态工作流。Meta 将发布 WildArtifactBench 中的 10 项任务。
Today we’re also previewing WildArtifactBench, an internal evaluation framework designed to assess agents on complex, real-world tasks across diverse deliverable formats. By using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth rubrics, it expands task coverage across practical multimodal workflows. We’re releasing 10 tasks from WildArtifactBench as a step forward in our ability to measure the real practical utility delivered by multimodal agents: research.meta.ai/wild-artifact-… 💬 9 🔄 7 ❤️ 80 👀 7131 📊 20 ⚡