模型精选

Datalab 发布 OmniExtractBench 基准测试,其 Agentic Plus 模型在测试中取得 93.94% 成绩

Recently, Datalab released OmniExtractBench, a benchmark that combines several benchmarks (our own E...

精选理由

Datalab 新发布的 OmniExtractBench 基准测试里,他们自己的 Agentic Plus 模型表现不错,值得看看。

Datalab 发布了 OmniExtractBench 基准测试,该测试结合了多个基准(包括 Datalab 自家的 ExtractBench 和 LongExtractBench)。他们使用 Agentic Plus 模式测试了该模型,通过修复基准测试中移除空值的兼容性问题后,在测试中取得了 93.94% 的成绩,与当前最高分持平或略高。

原文 · Jerry Liu

Recently, Datalab released OmniExtractBench, a benchmark that combines several benchmarks (our own E...

Recently, Datalab released OmniExtractBench, a benchmark that combines several benchmarks (our own ExtractBench, LongExtractBench, other vendors) to measure structured document quality. The benchmark measured our "Agentic Plus" extraction model, our most powerful mode for complex document extraction. We found a compatibility workaround in the benchmark’s integration that stripped null from fields that allowed it - removing a valid way to represent missing information. We restored that option while preserving the schema’s structure and required fields. With the same scorer, we were able to get to 93.94% on the benchmark. This puts it slightly higher (or at least equal) to the highest published scores. There's a lot of vendor benchmarks, but ultimately it's also important to evaluate over your own data. We'd be happy to help you get that set up. PR: github.com/datalab-to/omn… Our Agentic Plus mode for LlamaExtract is legitimately quite good, would love to you to try it out and let us know your thoughts! cloud.llamaindex.ai 💬 4 🔄 2 ❤️ 9 👀 947 📊 6 ⚡