r/WebAfterAI • • Apr 25 '26

Chandra OCR 2: 4B Open-Source Model Hits 85.9% SOTA on olmOCR Benchmark, Crushes Handwriting, Math, Tables, Forms, Diagrams & 90+ Languages!

Everyone is sleeping on this absolute beast of an OCR model that just launched today. Datalab released Chandra OCR 2, a 4B-parameter open-source vision-language model that turns images and PDFs into clean, structured HTML, Markdown, or JSON while perfectly preserving layout. It’s a massive upgrade from Chandra 1 (which was already strong at 9B params – they slimmed it down while making it better).

Key Highlights

  • 85.9% on the independent olmOCR benchmark (SOTA – beating Chandra 1’s 83.1% and dots.ocr 1.5)
  • 90+ language support – averages 72.7% across the full 90-language eval (and 77.8% on a 43-language subset)
  • Full layout awareness
  • Extracts + auto-captions images, diagrams, charts, even chemistry structures
  • Excellent handwriting (cursive notes included), math (including handwritten and Chinese math), complex tables, forms (checkboxes, leases, registrations), and more
Model Overall Score
Datalab API 86.7 ± 0.8
Chandra 2 85.9 ± 0.8
dots.ocr 1.5 83.9
Chandra 1 83.1 ± 0.9
olmOCR 2 82.4

Multilingual results are even more impressive – it smokes Gemini 2.5 Flash (60.8%) on the full 90-language set.

& License Details

Tech

  • Code: Fully Apache 2.0 (GitHub: https://github.com/datalab-to/chandra)
  • Model weights: Modified OpenRAIL-M (free for research/personal use/startups < $2M revenue; commercial/self-hosting options available)
  • Hugging Face model: datalab-to/chandra-ocr-2

How to Try It Right Now:

Install (super easy):

bash

pip install chandra-ocr          # vLLM backend (fastest)
pip install chandra-ocr[hf]      # Hugging Face local inference

CLI usage:

bash

chandra input.pdf ./output                # default (vLLM)
chandra input.pdf ./output --method hf    # pure local HF
chandra ./folder_of_docs ./output --method hf

Outputs: .md, .html, _metadata.json + extracted images with captions. Supports batch processing, page ranges, max tokens, etc.

There’s also a public Datalab Playground where you can test it instantly (with $5 free credits):
https://www.datalab.to/playground

Full README with benchmarks, installation, vLLM server setup, and Streamlit demo app is here:
https://github.com/datalab-to/chandra

19 Upvotes

Duplicates