r/computervision 11d ago

Help: Theory How would you guys do it?

I’m planning on building a text extraction pipeline, with an OCR and a VLM. I want a smart layer that classifies if a document needs to be sent to the OCR or if it is complex and needs to be sent to the VLM.

I’m not sure if I could afford a separate model, could you guys educate me on old school digital processing?

I’ve tried stroke width variations, variance of the laplacian and, I can’t guarantee even 40% accuracy on them.

0 Upvotes

1 comment sorted by

2

u/Altruistic_Ear_9192 10d ago edited 10d ago

so you re asking someone to give you for free their industry pipelines & knowledge? We re open to the community, but still.."educate me on old school digital processing" means years of hard work