Hugging Face和EleutherAI测了14款开源OCR,dots.mocr最准,97.6%准确率,每千页不到2美元,处理旧书够用了。
Hugging Face与EleutherAI的FineBooks项目测试了14款开源OCR模型,覆盖2000多页历史书籍。表现最好的dots.mocr字符准确率达97.6%,处理每千页成本低于2美元。团队认为该精度已可用于AI训练数据,但尚未达到学术转录标准。
Old OCR text cripples language model training, and FineBooks wants to fix that at scale
The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on more than 2,000 historical book pages. The top model, dots.mocr, hits 97.6 percent character accuracy at under two dollars per thousand pages. That's good enough for AI training data, but not yet for scholarly transcriptions, the team says. The article Old OCR text cripples language model training, and FineBooks wants to fix that at scale appeared first on The Decoder .