论文精选

Perplexity AI发布WANDR:智能体实体发现与事实验证评测基准

精选理由

Perplexity AI搞了个新评测WANDR,专门看智能体能不能找到一堆实体并核实每个的具体信息,比单一评分更清楚哪里不行。

WANDR是一个评测基准,用于测试智能体发现大规模实体集并验证每个实体特定事实的能力。该基准提供密集且可解释的评估信号,能揭示智能体在哪个环节失败。其流程还可作为半自动训练数据工厂。由Perplexity AI发布。

原文 · Perplexity

WANDR tests how well agents discover large sets of entities and verify specific facts about each one. It provides a dense, interpretable eval signal that reveals whether an agent fails, and where. The pipeline also doubles as a semi-automated factory for training data. 💬 1 🔄 1 ❤️ 5 👀 1487 📊 2 ⚡