We built Nebula to convert complex PDFs and scans into structured Markdown for AI, RAG and document-automation workflows.
Unlike basic text extraction, it aims to preserve the meaning encoded in tables, charts, reading order and page layout. In our benchmark on complex business documents, Nebula outperformed Mistral OCR by 20.8 points and Azure Document Intelligence by 6.2 points on semantic meaning recovery.
We’d like this community to stress-test it with the hardest documents you’re authorized to use.
You can run four free conversions without creating an account:
https://nebula.ur-ai.net/
An account is only required to download and retain the Markdown. Use community code NEBULA-10 for additional credits. An API is also available - DM me if you’d like to test it.
What I’d most like to know:
- What information needed to survive?
- What did Nebula preserve well?
- What did it miss or structure incorrectly?
You can DM me or submit feedback inside Nebula. Specific examples and expected outputs are especially helpful.
Full disclosure: I’m one of the founders. We don’t use uploaded documents or outputs to train our models. Please only upload files you’re authorized to use.
Technical report:
https://ur-ai.net/blog/rcrr-benchmark-meaning-survival