
Mistral AI is taking on document understanding with a new OCR model that doesn't just read text — it actually understands what it's looking at. Mistral OCR 4 extracts text, outputs bounding boxes, classifies elements like titles, tables, formulas, and signatures, and provides confidence scores at the word and page level. It supports PDF, DOC, PPT, and other common formats.
OCR 4 scores 85.20 on the multi-column and math formula benchmark OlmOCRBench, placing it first in the industry. It also achieved 93.07 on OmniDocBench. Early enterprise user Rogo tested it on dense chart financial QA and found that parsing accuracy matches cutting-edge agentic parsers, but at one-eighth the cost and one-seventeenth the latency.
The model already supports single-container self-hosting, so enterprises can run it on-premises for data privacy. Pricing starts at $4 per 1,000 pages for standard extraction, dropping to $2 with the batch interface. For $5 per 1,000 pages, the Document AI mode lets users pass a custom JSON schema, and the backend (e.g., mistral-small-2603) outputs formatted data directly.
Mistral OCR 4 is now available on Mistral Studio, Amazon SageMaker, and Microsoft Foundry.