Cohere新出的Parse 5能处理企业文档,比竞品在基准测试中表现更好,价格也透明。
Cohere推出Parse (parse-v5.0),这是一个23亿参数的视觉语言模型。该模型可将PDF、幻灯片和图像转换为包含HTML表格、边界框和图像描述的Markdown格式。Parse在ParseBench基准测试中得分为79.2,领先于Mistral OCR 4、Azure Document Intelligence和Databricks AI Parse。API定价为每1000页1.5美元,专用Model Vault实例起价为每月2500美元。
Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown
Cohere has released Parse (parse-v5.0), a 2.3B-parameter vision language model that converts PDFs, slides and images into Markdown with HTML tables, bounding boxes and image descriptions. It runs at $1.50 per 1,000 pages through the API, or on dedicated Model Vault instances from $2,500 a month. Cohere reports a ParseBench score of 79.2, ahead of Mistral OCR 4, Azure Document Intelligence and Databricks AI Parse — but that figure averages three of the benchmark's five dimensions and drops charts and visual grounding entirely. The post Cohere Releases Parse 5 (parse-v5.0): A 2.3B Vision Language Model That Turns Enterprise Documents Into Markdown appeared first on MarkTechPost .