论文精选

微软推出 CorpusMap:为文档集建实体地图,让智能体检索更准更省

Recommended. LLM agents love structure, so it's no surprise that a corpus improves agentic search.

精选理由

微软这篇论文很实用:给文档集预先建好实体地图,智能体检索准确率涨最多 11.7 分,token 还省一半。

微软团队发布 CorpusMap,针对智能体在大规模文档集上搜索的场景。它预先解析跨文档的重复实体,为每个实体生成一个链接到所有相关文档的页面,原始文档保持不变。智能体读文档时顺着实体跳转,避免重复搜索同一证据。在 7 个模型和 3 个基准上,答案质量提升 6.4 到 11.7 分,输入 token 减少 34% 到 57%,并超过 LLM Wiki 等 3 种导航层方案。

原文 · elvis

Recommended. LLM agents love structure, so it's no surprise that a corpus improves agentic search.

Recommended. LLM agents love structure, so it's no surprise that a corpus improves agentic search. DAIR.AI @dair_ai Banger paper from Microsoft and colleagues. If you run agents that search a large document collection, this one is worth your time. (bookmark it) They introduce CorpusMap, which resolves recurring entities across the collection in advance and gives each entity a page that links to every document that mentions it. The original documents stay in place. The agent reads a document, follows an entity to related documents, and avoids searching for the same evidence again. Across 7 models and three benchmarks, answer quality goes up 6.4 to 11.7 points while input tokens drop 34% to 57%. It also beats an LLM Wiki layer and three other navigation layers. The map can be built without LLM calls and updated as new documents arrive. Paper: arxiv.org/abs/2609.37226 Chat with Paper: academy.dair.ai/papers/follow-… 🔗 View Quoted Tweet 💬 4 🔄 0 ❤️ 8 👀 1473 📊 4 ⚡