论文

GeoNatureAgent:面向地理空间任务的工具型智能体评测基准发布

GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks

精选理由

一个专门测地理空间智能体的开源基准,103 个任务打九个模型,Claude Sonnet 4 第一但 DeepSeek V3.2 便宜 11 倍,做 Agent 评测的可以拿来就用。

GeoNatureAgent (GNA) 提供了一个固定十六工具的地理空间接口,以 MCP 服务器形式发布,评测时智能体是唯一变量。其旗舰实例包含 103 个任务(93 个主任务套件覆盖 18 个类别,另加 10 个对比扩展任务),基于西班牙和葡萄牙三个环境指标的开放地理空间 API。在三种 temperature-1.0 种子下评测九个 LLM:Claude Sonnet 4 以 61.7% 的全任务准确率居首,DeepSeek V3.2 以 57.9% 紧随其后,其余模型均不超过 53%。成本精度帕累托前沿主要由开源权重模型占据,DeepSeek V3.2 以约 11.3 倍更低的定价提供 Claude 约 93% 的能力。在严格全检查评分下,最佳模型比通用 GIS 基准报告的 85-97% 低 24-36 个百分点,但前四名的单检查部分得分(86-90%)相当,说明差距主要来自评分严格度而非任务难度。

原文 · arXiv: DeepSeek

GeoNatureAgent (GNA): A Framework and Benchmark for Pre-Production Evaluation of Tool-Using Agents on Geospatial and Environmental Tasks

Before tool-using LLM agents are deployed in environmental and geospatial workflows, teams need evidence that an agent reliably selects the right operations against real APIs. We introduce GeoNatureAgent (GNA), a framework for pre-production evaluation of tool-using agents: a fixed sixteen-tool geospatial interface published as a Model Context Protocol (MCP) server, so the agent under test is the only variable, scored against an identical tool layer, task suite, and deterministic scorer. Its flagship instance is a 103-task benchmark (a 93-task main suite across 18 categories plus a ten-task comparison expansion) evaluated against an open, self-hostable geospatial API serving three environmental indicators across Spain and Portugal. We evaluate nine LLMs under three temperature-1.0 seeds, reporting capability and per-case cost as orthogonal axes. (1) Claude Sonnet 4 achieves the highest capability (61.7% +/- 0.7% on all 103 tasks; 60.8% on the main suite), followed closely by DeepSeek V3.2 (57.9%), while no other model exceeds 53%; (2) the cost-accuracy Pareto frontier is mostly open-weight, with DeepSeek V3.2 offering 93% of Claude's capability at 11.3x lower list-price cost; (3) under strict all-checks scoring the best model sits 24-36 points below the 85-97% reported on general-purpose GIS benchmarks, whereas per-check partial credit for the top four models (86-90%) is comparable, so much of that gap reflects scoring strictness rather than task difficulty alone. The MCP server, evaluation harness, benchmark, and API are publicly available; swapping the tool executors and task suite instantiates an equivalent benchmark for any geospatial domain.