论文精选

微软发布 CorpusMap:让智能体跨文档检索更省 token

精选理由

微软发了篇论文叫 CorpusMap,提前把文档集合里的实体建好索引页,智能体检索时准确率涨最多 11.7 分,token 还省一半多。

微软及合作者发表论文 CorpusMap,面向运行大规模文档集合检索的智能体。该方法提前解析集合中的重复实体,为每个实体建立页面并链接到所有提及它的文档,原始文档保持不变。智能体读取文档后可沿实体跳转到相关文档,避免重复搜索同一证据。在 7 个模型和三个基准上,答案质量提升 6.4 至 11.7 分,输入 token 减少 34% 至 57%,并超过 LLM Wiki 层和另外三种导航层。地图构建无需 LLM 调用,新文档到达时可增量更新。

原文 · DAIR.AI

Banger paper from Microsoft and colleagues.

If you run agents that search a large document collection, this one is worth your time.

(bookmark it)

They introduce CorpusMap, which resolves recurring entities across the collection in advance and gives each entity a page that links to every document that mentions it.

The original documents stay in place.

The agent reads a document, follows an entity to related documents, and avoids searching for the same evidence again.

Across 7 models and three benchmarks, answer quality goes up 6.4 to 11.7 points while input tokens drop 34% to 57%.

It also beats an LLM Wiki layer and three other navigation layers.

The map can be built without LLM calls and updated as new documents arrive.

Paper: https://t.co/GimGee8u8V

Chat with Paper: https://t.co/0GOcC169at