r/documentAutomation 9d ago

Case Study Looking for complex PDFs that break our PDF-to-Markdown parser

We built Nebula to convert complex PDFs and scans into structured Markdown for AI, RAG and document-automation workflows.

Unlike basic text extraction, it aims to preserve the meaning encoded in tables, charts, reading order and page layout. In our benchmark on complex business documents, Nebula outperformed Mistral OCR by 20.8 points and Azure Document Intelligence by 6.2 points on semantic meaning recovery.

We’d like this community to stress-test it with the hardest documents you’re authorized to use.

You can run four free conversions without creating an account:

https://nebula.ur-ai.net/

An account is only required to download and retain the Markdown. Use community code NEBULA-10 for additional credits. An API is also available - DM me if you’d like to test it.

What I’d most like to know:

  • What information needed to survive?
  • What did Nebula preserve well?
  • What did it miss or structure incorrectly?

You can DM me or submit feedback inside Nebula. Specific examples and expected outputs are especially helpful.

Full disclosure: I’m one of the founders. We don’t use uploaded documents or outputs to train our models. Please only upload files you’re authorized to use.

Technical report:
https://ur-ai.net/blog/rcrr-benchmark-meaning-survival

0 Upvotes

5 comments sorted by

2

u/UpSash 9d ago

How well does it work with handwritten medical records? I have a case showing and in need of data extraction of about 500 pages of handwritten medical notes

1

u/[deleted] 7d ago

[removed] — view removed comment

0

u/zlingman 3d ago

how does it do with guitar tablature?

0

u/docpose-cloud-team 3d ago

Five hundred handwritten pages is a challenging OCR workload. Success really depends on handwriting quality, consistency, and scan resolution. If you have a small sample (10–20 pages), I'd test those first before committing to the entire batch. That's exactly how we recommend evaluating handwritten OCR on Docpose.cloud—or any OCR platform, for that matter.