WorldBench:多语言智能体文化基准测试

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

精选理由

研究人员发布WorldBench基准,测试多语言智能体在真实场景中的表现,发现当前模型在长周期任务中表现脆弱。

AI 摘要

WorldBench是一个包含1600个任务的多语言基准测试,覆盖7种语言和8种文化。该基准测试通过结构化动作让智能体在沙盒环境中执行日常任务。研究显示前沿模型仅达到49.2%的约束任务成功率(CTS),所有模型在正确性和环境保持方面存在显著差距。

原文 · arXiv cs.AI

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to realistic, grounded scenarios. To address these concerns, we present WorldBench: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions. WorldBench comprises 1,600 tasks across seven languages and eight cultures, filtered and refined through feedback from human annotators with language- and culture-specific expertise. For evaluation, we extend metrics from previous works and introduce Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and other complementary metrics through deterministic and LLM-as-a-Judge evaluations. Our experiments show that frontier models reach only 49.2% CTS, with all models demonstrating large gaps between correctness and environment preservation. We thereby show that current agents remain brittle in multilingual, agentic scenarios, especially for long-horizon tasks and under state-preservation constraints