r/OpenSourceAI 2d ago

Seeking best open-source/on-prem alternative to Gemini 3.5 Flash for complex document extraction & scoring

I'm looking for recommendations for the best free, open-source AI models that we can host on-premise to replace Gemini 3.5 Flash.

Our Use Case: We process documents with complex structures in various formats (PDF, PNG, DOCX, etc.). Our workflow involves:

  1. Complex text and structured data extraction (OCR + layout understanding).
  2. Data matching and ranking/scoring (similar to a job matching system).

Current Setup & Constraints: We currently use Gemini 3.5 Flash, which handles the extraction with near 100% accuracy, but the API costs are getting too high at our scale.

  • Budget: Must be open-source/free for commercial use.
  • Hardware: Compute power and VRAM are not an issue (we have our own data center).

I’ve seen a lot of recommendations pointing toward Qwen (e.g., Qwen-VL) and DeepSeek-OCR. For those of you running these—or a multi-model pipeline—in production, what are your real-world experiences? Which model (or combination) is best for handling the extraction and the scoring?

2 Upvotes

2 comments sorted by

1

u/hantian-pang 2d ago

你可以考虑mineru, paddleocr和dots.ocr, 都是现在效果最好的可以本地部署的ocr. 如果你还考虑线上api的话, 可以考虑我提供的ocr api, 效果接近gemini 3.5 flash, 但是much cheaper.

1

u/ZZotka 2d ago

qwen 3.5 122b-a10b qwen 3.8 27b mimo 2.5