Artificial intelligence

Cohere Launches Parse 5 for AI Document Processing at $1.50 per 1,000 Pages

Cohere has launched Parse 5, a 2.3-billion-parameter vision-language model that converts PDFs, presentations, and images into structured Markdown, with support for tables, image descriptions, and element localization. The model does not top the accuracy comparison published by the company, but it is positioned to reduce the cost of large-scale processing, priced at $1.50 per 1,000 pages.

2026-08-28
4 min read
6 views
فريق تحرير certi.news
Cohere Launches Parse 5 for AI Document Processing at $1.50 per 1,000 Pages

Cohere announced the availability of Parse 5 for processing PDFs, presentations, and scanned images and converting them into structured Markdown, targeting organizations that need to feed large volumes of documents into AI applications and software agents. The model is available through the Cohere API, Model Vault, Microsoft Foundry, and AWS SageMaker.

The company positions Parse 5 differently from larger general-purpose AI models: not as the most accurate option, but as an attempt to balance accuracy and cost when processing millions of pages. Cohere set API usage pricing at $1.50 per 1,000 pages, while Model Vault enables managed deployment on a secure single-tenant platform for larger volumes.

One Model Instead of a Separate Processing Pipeline

Parse 5 is based on a vision-language model with 2.3 billion parameters and built on Cohere Labs’ North-Micro-Vision-Instruct architecture. Its context window is 8,192 tokens, while the model size is estimated at approximately 4.6 gigabytes. The model accepts a PDF, PowerPoint, or JPEG page as a base64-encoded image, then returns its contents in reading order as Markdown.

Tables are converted to HTML, while the model also adds image descriptions and bounding-box coordinates for tables and images. The default mode returns a Markdown string for each page, whereas blocks mode provides typed elements, with each table containing its own HTML, description, and coordinates. Cohere believes this format helps track the source of information at the citation level within agent applications.

Arabic, English, French, German, Italian, Japanese, Korean, Portuguese, and Spanish each achieve stable accuracy according to the source, with lower-accuracy zero-shot support for other languages.

Lower Accuracy Than Larger Models, With a Different Price Point

According to the ParseBench comparison published by Cohere, Parse 5 scored 79.2 across three dimensions: tables, content fidelity, and semantic formatting. It was surpassed by GPT-5.5 models with a score of 84.4, Opus 4.8 with 84.3, and Gemini 3.5 Flash with 81.8. In contrast, Parse 5 outperformed LlamaParse’s Cost Effective category, which scored 78.3, as well as Mistral OCR 4 with 74.5, Databricks AI Parse with 72.4, and Azure Document Intelligence with 69.3.

The comparison excludes the Layout and Chart dimensions, and Cohere says this is due to the product’s scope. The model returns text in reading order rather than providing coordinates for every text element, and it describes charts instead of extracting their underlying data. The company says chart-data extraction is planned for a future release.

What Changes Practically for Organizations?

Parse 5’s primary value lies in reducing reliance on a large general-purpose model for every page, rather than in absolute superiority over processing tools. Cohere estimated that, in a typical workflow at a financial-services company processing 750 million documents annually, using Parse 5 instead of GPT-5.5 could reduce costs by more than 98%. However, this percentage is the company’s estimate for a modeled workflow, not the result of an audited deployment or an independent study.

This comparison confirms that choosing a document-analysis tool should not depend on a single benchmark score. If tables, headings, or reading order are lost during ingestion, embedding models or larger models later in the pipeline will not be able to recover the missing structure. Organizations will therefore need to test Parse 5 on their most difficult documents and measure retrieval accuracy and end-task performance, rather than relying solely on an assessment of the appearance of the extracted text.

The expert opinions cited in the source support this cautious use; they view Parse 5 as falling between traditional OCR and running a high-cost general-purpose model on every page. The open question remains whether the provision of structure and element coordinates, together with the low price, will actually lead to better outcomes in end applications, particularly for documents rich in charts or complex layouts.

News source
VentureBeat Startups & Funding
Open original source ↗
ف
Author

فريق تحرير certi.news

In the same category

You may also like

View all news