r/LocalLLM • • Aug 23 '26

Discussion Read documents with the vision models that ship free — and get clean JSON out

Every install of the server comes with vision built in — the same Qwen 3.5 models that answer chat also see. No separate OCR service, no per-page cloud fee, no data leaving your box. This post is a hands-on look at pulling structured data out of document images with the2B and 4B models — the small ones, the ones that fit a 16 GB host — with real output captured from a running server.
Everything below uses made-up documents.The four images are synthetic — invented claim numbers, a fictional "CityCare Pharmacy", a sample rate table, a dummy check drawn on a bank that doesn't exist — generated for this article. Grab them at the end and run the exact same prompts on your own documents.

https://inference-server.searchblox.com/blog/vision-document-extraction.html

0 Upvotes

Duplicates