Firecrawl把自己最快的PDF解析引擎开源了,200个PDF只要2.8秒,还能提取表格图表,比传统OCR省事多了。
Firecrawl开源了PDF解析引擎pdf-inspector,该引擎用Rust编写,可在约20ms内完成PDF分类。它提取内容的速度为每页0.002秒,测试中处理200个PDF仅耗时2.8秒。pdf-inspector能高质量提取表格和图表,并生成干净的Markdown格式。项目代码已托管在GitHub上。
we open sourced the fastest pdf parser engine pdf-inspector powers /parse together with our custom ...
we open sourced the fastest pdf parser engine pdf-inspector powers /parse together with our custom OCR models 0.002s per page Nicolas Camara @nickscamara_ we built pdf-inspector so agents can process PDFs without waiting on OCR. it classifies any PDF in ~20ms and extracts clean markdown locally → 200 PDFs processed in 2.8s → top quality in extracting tables + graphs → built in rust → open source github.com/firecrawl/pdf-… 🔗 View Quoted Tweet 💬 1 🔄 4 ❤️ 62 👀 4549 📊 15 ⚡
- Geek07-31 15:32原文