LlamaIndex给LlamaParse加了新模式,专门抽长文档里的海量字段,准确率比Claude Code和Codex高10-20%,每个字段还带出处。
LlamaIndex发布LlamaExtract Agentic Plus模式,专为50页以上长文档设计,可提取10k-100k个字段,准确率超过94%。每个提取字段附带置信度分数和边界框,便于溯源。在ExtractBench基准测试中,该模式比Claude Code Opus 4.8和Codex GPT-5.6等通用编码智能体准确率高10-20%。测试案例包括一份FTX债权人矩阵,含75k个字段、共114页。
We tuned an AI agent that can do large-scale document extraction from long docs (50+ pages, some wit...
We tuned an AI agent that can do large-scale document extraction from long docs (50+ pages, some with 10k-100k fields) with 94%+ accuracy 📈 It uses a harness + model set that is tuned specifically for reasoning over extracting out complex information from complex docs. Each extracted field comes with a confidence score as well as a bounding box denoting where it came from. It does 10-20% better in accuracy than generalized coding agent harnesses (e.g. Claude Code Opus 4.8 and Codex GPT-5.6). Check out the video below for a demo. The mode is called LlamaExtract Agentic Plus. Learn more about our extraction benchmark and agentic plus mode here: llamaindex.ai/blog/introduci… 7 If you have very complex document extraction needs, come check out LlamaParse! cloud.llamaindex.ai 8 Your browser does not support the video tag. 🔗 View on Twitter Jerry Liu @jerryjliu0 Our "agentic plus" extractor in LlamaParse is great for extracting out massive volumes of fields (e.g. 10k-100k+ fields) from long documents (100-500 pages) We tested with a doc that contains a giant matrix of all creditors for FTX 💸 (75k fields, 114 pages). See screenshot below. You can find full benchmark results on ExtractBench extractbench.ai r). In the meantime, come try out the mode in our playground! cloud.llamaindex.ai A 🔗 View Quoted Tweet 💬 4 🔄 2 ❤️ 7 👀 1375 📊 5 ⚡