论文精选73°

CorporateBench: 大规模问答基准测试

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

精选理由

CorporateBench 填补了企业级大模型评估空白,用23万文档测试了5个模型在真实规模下的表现。

AI 摘要

CorporateBench 是一个人工验证的多任务问答基准,评估语料超过23万份文档。该基准通过四个模拟企业(员工规模从12到10,000人)评估大模型在信息提取和知识库查询两个维度的表现。研究显示,随着输入规模接近真实场景,大模型性能显著下降。

原文 · arXiv cs.AI

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't want to share internal communications, and synthetic datasets have been overly simple. We present CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, with evaluation corpora surpassing 230,000 documents. CB evaluates LLMs across two dimensions (information extraction and knowledge base querying) through four synthetically generated firms ranging from 12 to 10,000 employees. Each corpus is sampled from a temporally evolving knowledge base describing a consistent world, guaranteeing cross-document logical consistency even across hundreds of thousands of documents. We evaluate five LLMs on CB, revealing increasingly poor performance as input size approaches realistic scales. CB provides LLM developers a metric for corporate communication reasoning, filling a crucial gap in the benchmarking ecosystem.