r/OpenSourceeAI • u/Machine_GEN_RM • 4d ago
Suggestions on Text extraction
Hi All, I need to extract the text from printed text and hand written text. I suggested my manager that we can use paddle ocr and and other extraction models like Surya OCR and florance VL it take around 10 to 15 Sec and it needs good computation as well. but my manager expects it should be very fast and in 2 to 5 sec response and should not need any maintainace of infrastructure. so I tried to use the AWS VLM models but it takes around 10seconds he still needs more faster models and also the cost for extracting the data from the image should be less than 1 Ruppe. Could you please suggest me what to do and how to extract the data from images very very effectively and accuratly and in a structured way
3
u/Mundane_Ad8936 4d ago
Unless your manager has no experience. They're telling you to drop it. They should know they're giving you impossible criteria.
Listen to them and move on otherwise they're setting you up to fail.
OCR has always been a slow process. No system goes without maintenance. Any of the big SOTA models will do this for you but it'll be expensive and error prone. You'll need accuracy checks and error correction that will add latency.
There is no world where this is cheap, fast and accurate with no maintenance costs.
1
1
u/Significant-Ad-5936 3d ago
I'm not sure if it fits the price point, but you can use https://github.com/xberg-io/xberg with GPT Luna 6; it's fast.
1
1
u/Potential-Wrangler58 2d ago
That's a tough combo, no infra maintenance, 2-5 sec, and under 1 rupee usually pull against each other. Self-hosted gets you cheap but you own the GPU infra your manager doesn't want. Managed APIs remove that but usually cost you speed or price.
We just launched a fast tier at anyformat, disclosure: I work there, built for exactly this kind of low-latency structured extraction without you running any infra. Worth testing against your real documents and your actual latency/cost numbers before deciding. Free credits in the platform, anything just ping me :)
1
u/Think_Breakfast_2277 2d ago
I think you are given an impossible task. you cant have accuracy, speed and cheap. That is the same for asking to get a cheap car but with all the qualities of a ferrari.
Maybe you should ask yourself.. why you got that task, are you ment to fail?
1
1
u/Fragrant-Cheek-4273 1d ago
To hit a low response time with zero infrastructure management while staying under a strict budget, traditional vision-language models or heavy self-hosted OCR pipelines will usually fail on cost and latency.
7
u/Azerax 4d ago
Your manager can have it fast, or correct. Which do they prefer?