AgentHop 基准发布:1,011 道题拆解智能体失败原因
AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering
一个把智能体准确率拆开看的新基准,能看出各家模型是乱搜还是不敢搜,做 agent 评估的可以拿去用。
论文推出诊断型基准 AgentHop,包含 1,011 道多选题和七工具沙盒,在固定 token、轮次、工具调用约束下测试。它把单一准确率拆成检索、综合、工具调用、资源管理四个维度,覆盖 19 个模型。结果显示行为按模型家族聚集:GPT 提前行动,Anthropic 和 GLM 先验证再行动,DeepSeek 和 Kimi 过度搜索,Gemini-3 Pro 较均衡。同家族内部也有分化,Claude Opus 4.6 检索更多,Sonnet 4.6 综合更好。
AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering
Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagnose and improve agentic systems. Existing benchmarks, however, tend to focus on a single leaderboard score, leaving the underlying failure modes opaque. To fill this gap, we introduce AgentHop, a diagnostic benchmark of 1,011 multiple-choice questions paired with a controlled seven-tool sandbox under fixed token, turn, and tool-call constraints. AgentHop reveals model vulnerabilities by dissecting a single accuracy score along four axes of agent operation: retrieval, synthesis, tool-call, and resource management. Across 19 models, we find that behavior clusters by model family, with tool-call signatures revealing distinct family fingerprints: GPT models commit early, Anthropic and GLM checkpoints verify before committing, DeepSeek and Kimi over-search, and Gemini-3 Pro stays balanced. Decomposed axes further expose within-family structure: Claude Opus 4.6 and Sonnet 4.6 land within one accuracy point yet diverge on retrieval-versus-synthesis emphasis, with Opus retrieving more and Sonnet synthesizing better. We release the full benchmark set and the harness to support diagnostic agent benchmarking.