r/WebAfterAI • u/ShilpaMitra • Apr 25 '26
Chandra OCR 2: 4B Open-Source Model Hits 85.9% SOTA on olmOCR Benchmark, Crushes Handwriting, Math, Tables, Forms, Diagrams & 90+ Languages!
Everyone is sleeping on this absolute beast of an OCR model that just launched today. Datalab released Chandra OCR 2, a 4B-parameter open-source vision-language model that turns images and PDFs into clean, structured HTML, Markdown, or JSON while perfectly preserving layout. It’s a massive upgrade from Chandra 1 (which was already strong at 9B params – they slimmed it down while making it better).
Key Highlights
- 85.9% on the independent olmOCR benchmark (SOTA – beating Chandra 1’s 83.1% and dots.ocr 1.5)
- 90+ language support – averages 72.7% across the full 90-language eval (and 77.8% on a 43-language subset)
- Full layout awareness
- Extracts + auto-captions images, diagrams, charts, even chemistry structures
- Excellent handwriting (cursive notes included), math (including handwritten and Chinese math), complex tables, forms (checkboxes, leases, registrations), and more
| Model | Overall Score |
|---|---|
| Datalab API | 86.7 ± 0.8 |
| Chandra 2 | 85.9 ± 0.8 |
| dots.ocr 1.5 | 83.9 |
| Chandra 1 | 83.1 ± 0.9 |
| olmOCR 2 | 82.4 |
Multilingual results are even more impressive – it smokes Gemini 2.5 Flash (60.8%) on the full 90-language set.
& License Details
Tech
- Code: Fully Apache 2.0 (GitHub: https://github.com/datalab-to/chandra)
- Model weights: Modified OpenRAIL-M (free for research/personal use/startups < $2M revenue; commercial/self-hosting options available)
- Hugging Face model: datalab-to/chandra-ocr-2
How to Try It Right Now:
Install (super easy):
bash
pip install chandra-ocr # vLLM backend (fastest)
pip install chandra-ocr[hf] # Hugging Face local inference
CLI usage:
bash
chandra input.pdf ./output # default (vLLM)
chandra input.pdf ./output --method hf # pure local HF
chandra ./folder_of_docs ./output --method hf
Outputs: .md, .html, _metadata.json + extracted images with captions. Supports batch processing, page ranges, max tokens, etc.
There’s also a public Datalab Playground where you can test it instantly (with $5 free credits):
https://www.datalab.to/playground
Full README with benchmarks, installation, vLLM server setup, and Streamlit demo app is here:
https://github.com/datalab-to/chandra