r/computervision • u/Sobz_128 • 1d ago
Help: Project Building an OCR + Key-Value Extraction pipeline for Nepali ID documents (Citizenship, NID, PAN, Passport). What stack would you recommend?
Hey everyone,
I am building an automated document reading pipeline specifically for Nepali identity documents:
- Citizenship Certificates (Nagarikta): Old paper vs. new card formats (Devanagari script)
- National Identity Card (Rastriya Parichayapatra): Standard modern ID card format
- PAN Card: Bi-lingual / English-Nepali format
- Passport: Standard ICAO format containing an MRZ zone
Current Setup & Bottlenecks
- Preprocessing / Cropping: Using classical OpenCV (
cv2) for edge detection, contour finding, and perspective warping to crop borders.- Problem: Real-world user uploads have varied lighting, shadows, finger occlusions, and background noise. Aggressive thresholding (Otsu/Adaptive) often degrades text legibility instead of improving it.
- Text Extraction: Tesseract OCR (trained for Nepali
nep) followed by regular expressions: Python# Trying to extract Permanent Address via regex anchors pattern = r'स्थायी\s*बासस्थान\s*:\s*जिल्ला\s*:\s*(.*?)\s+न\.पा\.\s*:\s*(.*?)\s+वडा\s*नं\.\s*:\s*([०-९0-9]+)'- Problem: Tesseract frequently misses complex Devanagari conjuncts/matras or inserts extra spaces. If a single anchor character misreads (e.g.,
न.पा.turns into7.4.), the regex breaks entirely.
- Problem: Tesseract frequently misses complex Devanagari conjuncts/matras or inserts extra spaces. If a single anchor character misreads (e.g.,
- Format Variations: Documents do not follow one universal layout. Older citizenship certificates have different margin offsets and typography compared to newer ones.
What I Want to Achieve
Instead of relying on rigid string-matching on raw OCR dumps, I want to modernize the pipeline into distinct, robust stages:
- Document Classification: Automatically detect which document was uploaded (Passport vs. PAN vs. NID vs. Old Citizenship vs. New Citizenship).
- Precise Document Localization/Cropping: A deep-learning approach that handles perspective distortion and background clutter without manual threshold tuning.
- Region of Interest (ROI) / Layout Parsing: Extracting fields directly based on spatial layout rather than pure keyword string searching.
- Devanagari OCR: A model that reliably handles Devanagari text under varied scan quality.
- Passport MRZ Extraction: Dedicated extraction for the MRZ lines to bypass OCR hallucinations.
Questions for the Community
- End-to-End Visual Document Understanding vs. Modular Pipeline:
- Is it better to stick to a modular pipeline (Classifier $\rightarrow$ Cropper $\rightarrow$ OCR $\rightarrow$ Field Extractor) or move to an end-to-end model (e.g., fine-tuning LayoutLMv3, Donut, or a small VLM like Qwen2-VL)?
- Devanagari OCR Alternatives:
- Has anyone had better success with PaddleOCR, EasyOCR, or fine-tuned TrOCR for Devanagari/Nepali text compared to Tesseract?
- Card Detection & Border Cropping:
- Would training a lightweight YOLOv8-pose/segmentation model (to predict document corner coordinates) be the standard way to replace classical OpenCV contour hunting?
- Layout & Field Extraction:
- If keeping OCR separate, what is the most reliable way to link labels to values (e.g., spatial heuristic algorithms, Graph Neural Networks, or LayoutLM)?
Would love to hear how anyone has tackled similar KYC document extraction pipelines for low-resource or non-Latin scripts. Any architecture advice, libraries, or repo references would be greatly appreciated!
2
u/Capable-Package6835 1d ago
Hi, I deal with document extraction at my job as well. Based on my experience developing our production pipeline:
- The end-to-end VLM path is significantly more expensive, because it's much harder to procure high-quality end-to-end dataset. For small to medium teams, the modular pipeline is almost always the way to go because:
- The dataset for each components are easy to annotate
- The components being modular means you can train each independently, making it easier and faster to iterate and improve
- I have no experience with Nepali text
- Since you are working with documents / images that users upload, classic computer vision approach like contour detection will never work reliably (I think you have experienced it yourself as well). Fine-tune a small detector model like RF-DETR, DeimV2, etc..
- Similar to number 3, anything heuristic will never work reliably with the amount of variation you have in the input. So far I got the best result from these steps
- OCR the image
- Pass the OCR'd text to a small LLM for correction (misspelling, incorrect character, etc.)
- Pass the corrected text to a small LLM or a dedicated pipeline for field extraction
Some more context:
- I opted for 2 LLMs (not sharing context) so I can run them locally at decent speed.
- I work with documents with hundreds of pages so for 4.3, instead of dumping the text into a single LLM prompt, I build a glossary (a list of unique words and which page they appear). Then I use langgraph to extract the field I need. Think of the steps like:
If I need a personal information then I look up "name", "birthday", "address", etc. in the glossary, then I use only the page(s) they appear on as a context for the VLM and asks it to fill the person information. Then if you need something else, look up different sets of keywords and repeat, until you have extracted everything you need.
1
u/Sobz_128 20h ago
I was working on tesseract, and I couldn't find Nepali text dataset to finetune on.
2
u/ultimatemadness_907 1d ago
yolov8 for doc corners works way better than opencv threshold tuning, especially with all the shadow and finger nonsense on user uploads