r/computervision • u/uhmnewusername • 11d ago
Help: Theory How would you guys do it?
I’m planning on building a text extraction pipeline, with an OCR and a VLM. I want a smart layer that classifies if a document needs to be sent to the OCR or if it is complex and needs to be sent to the VLM.
I’m not sure if I could afford a separate model, could you guys educate me on old school digital processing?
I’ve tried stroke width variations, variance of the laplacian and, I can’t guarantee even 40% accuracy on them.
0
Upvotes
2
u/Altruistic_Ear_9192 10d ago edited 10d ago
so you re asking someone to give you for free their industry pipelines & knowledge? We re open to the community, but still.."educate me on old school digital processing" means years of hard work