想知道数据智能体怎么选?DataSpace测了6个模型和5个框架,换框架能让准确率差15个点,值得看。
新研究发布DataSpace基准,用于评估数据智能体在异构数据工作区生成可验证表格结果的能力。该基准包含410个跨语言任务,覆盖7439个工件,总数据量15.01GB,涉及CSV、JSON、SQLite、Markdown、PDF和视频等格式。在六个前沿多模态模型和五个常用智能体框架的测试中,最佳准确率为66.34%。固定模型主干时,更换框架可使准确率变动15.36个百分点。多模态证据整合和连接操作在六个骨干模型上均降低了准确率。
Harness choice is a big deal. So much room to advance and improve results across the board with age...
Harness choice is a big deal. So much room to advance and improve results across the board with agent harnesses. Great paper highlighting this. New research releases DataSpace, a benchmark where data agents produce verifiable tabular results from heterogeneous workspaces. 410 cross-language tasks over 7,439 artifacts totalling 15.01 GB across CSV, JSON, SQLite, Markdown, PDF, and video. Across six recent frontier multimodal models and five widely used agent harnesses, the best accuracy reaches 66.34%. With the backbone held fixed, swapping the harness moves accuracy by 15.36 points. Multimodal evidence integration and joins reduce accuracy across all six backbones. The benchmark is nowhere near saturated. Paper: arxiv.org/abs/2608.03451 Track more trending AI papers in our academy: academy.dair.ai 💬 4 🔄 7 ❤️ 33 👀 4272 📊 15 ⚡