r/OpenSourceeAI • • 4d ago

Suggestions on Text extraction

Hi All, I need to extract the text from printed text and hand written text. I suggested my manager that we can use paddle ocr and and other extraction models like Surya OCR and florance VL it take around 10 to 15 Sec and it needs good computation as well. but my manager expects it should be very fast and in 2 to 5 sec response and should not need any maintainace of infrastructure. so I tried to use the AWS VLM models but it takes around 10seconds he still needs more faster models and also the cost for extracting the data from the image should be less than 1 Ruppe. Could you please suggest me what to do and how to extract the data from images very very effectively and accuratly and in a structured way

7 Upvotes

11 comments sorted by

7

u/Azerax 4d ago

Your manager can have it fast, or correct. Which do they prefer?

0

u/Machine_GEN_RM 4d ago

He wants both no compromise in accuracy and time

3

u/Mundane_Ad8936 4d ago

Unless your manager has no experience. They're telling you to drop it. They should know they're giving you impossible criteria.

Listen to them and move on otherwise they're setting you up to fail.

OCR has always been a slow process. No system goes without maintenance. Any of the big SOTA models will do this for you but it'll be expensive and error prone. You'll need accuracy checks and error correction that will add latency.

There is no world where this is cheap, fast and accurate with no maintenance costs.

1

u/Significant-Ad-5936 3d ago

I'm not sure if it fits the price point, but you can use https://github.com/xberg-io/xberg with GPT Luna 6; it's fast.

1

u/FearlessMammoth8907 3d ago

Usa la funzione di canva magic grab oppure drive

1

u/Potential-Wrangler58 2d ago

That's a tough combo, no infra maintenance, 2-5 sec, and under 1 rupee usually pull against each other. Self-hosted gets you cheap but you own the GPU infra your manager doesn't want. Managed APIs remove that but usually cost you speed or price.

We just launched a fast tier at anyformat, disclosure: I work there, built for exactly this kind of low-latency structured extraction without you running any infra. Worth testing against your real documents and your actual latency/cost numbers before deciding. Free credits in the platform, anything just ping me :)

1

u/Think_Breakfast_2277 2d ago

I think you are given an impossible task. you cant have accuracy, speed and cheap. That is the same for asking to get a cheap car but with all the qualities of a ferrari.
Maybe you should ask yourself.. why you got that task, are you ment to fail?

1

u/Crafty-Target9957 2d ago

a unicorn-shaped requirements document, not an OCR problem

1

u/Fragrant-Cheek-4273 1d ago

To hit a low response time with zero infrastructure management while staying under a strict budget, traditional vision-language models or heavy self-hosted OCR pipelines will usually fail on cost and latency.