研究:小型本地模型做信息抽取的精度与能耗权衡
The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models
本地跑小模型抽文档信息?这篇把能耗账算清了:批处理省 38-85% 电,纯文本场景用文本模型比视觉模型又准又省。
arXiv 论文 2609.31341 研究在本地部署 ≤8B 参数模型处理隐私敏感文档时,精度与能耗的取舍。在 Kleister-NDA 合同和 VRDU 表单两个基准上测试发现,批处理是最大的节能杠杆,可将每页能耗降低 38-85% 且不损失精度。神经 OCR 每页能耗是传统 OCR 的 17 倍,且始终无法进入帕累托前沿。文档类型决定最优方案:版面丰富的文档用视觉语言模型,接近纯文本的文档用小型纯文本模型配合廉价解析器。
The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ parameter) text-only and vision--language models, evaluated on both accuracy and energy over a design space spanning input representation, model family, and inference configuration. Benchmarking on the near-plain-text Kleister-NDA contracts and the layout-rich VRDU forms, we find that batching is the dominant energy lever, cutting energy per page by 38-85% at no cost in accuracy, while FP8 quantization saves 27-32% when requests are served one at a time but less than 1mWh per page (9-19%) once batching is applied. Preprocessing dominates what remains: neural OCR costs $17\times$ more energy per page than classical OCR and never reaches the Pareto frontier. Which representation wins flips with the type of document: vision--language models on layout-rich documents and small text-only models with a cheap parser on near-plain text, where they are both more accurate and cheaper than any vision--language configuration. Our work yields concrete guidelines for energy-efficient, privacy-compliant local information extraction.