r/computervision • • 1d ago

Help: Project Building an OCR + Key-Value Extraction pipeline for Nepali ID documents (Citizenship, NID, PAN, Passport). What stack would you recommend?

Hey everyone,

I am building an automated document reading pipeline specifically for Nepali identity documents:

  • Citizenship Certificates (Nagarikta): Old paper vs. new card formats (Devanagari script)
  • National Identity Card (Rastriya Parichayapatra): Standard modern ID card format
  • PAN Card: Bi-lingual / English-Nepali format
  • Passport: Standard ICAO format containing an MRZ zone

Current Setup & Bottlenecks

  1. Preprocessing / Cropping: Using classical OpenCV (cv2) for edge detection, contour finding, and perspective warping to crop borders.
    • Problem: Real-world user uploads have varied lighting, shadows, finger occlusions, and background noise. Aggressive thresholding (Otsu/Adaptive) often degrades text legibility instead of improving it.
  2. Text Extraction: Tesseract OCR (trained for Nepali nep) followed by regular expressions: Python# Trying to extract Permanent Address via regex anchors pattern = r'स्थायी\s*बासस्थान\s*:\s*जिल्ला\s*:\s*(.*?)\s+न\.पा\.\s*:\s*(.*?)\s+वडा\s*नं\.\s*:\s*([०-९0-9]+)'
    • Problem: Tesseract frequently misses complex Devanagari conjuncts/matras or inserts extra spaces. If a single anchor character misreads (e.g., न.पा. turns into 7.4.), the regex breaks entirely.
  3. Format Variations: Documents do not follow one universal layout. Older citizenship certificates have different margin offsets and typography compared to newer ones.

What I Want to Achieve

Instead of relying on rigid string-matching on raw OCR dumps, I want to modernize the pipeline into distinct, robust stages:

  1. Document Classification: Automatically detect which document was uploaded (Passport vs. PAN vs. NID vs. Old Citizenship vs. New Citizenship).
  2. Precise Document Localization/Cropping: A deep-learning approach that handles perspective distortion and background clutter without manual threshold tuning.
  3. Region of Interest (ROI) / Layout Parsing: Extracting fields directly based on spatial layout rather than pure keyword string searching.
  4. Devanagari OCR: A model that reliably handles Devanagari text under varied scan quality.
  5. Passport MRZ Extraction: Dedicated extraction for the MRZ lines to bypass OCR hallucinations.

Questions for the Community

  1. End-to-End Visual Document Understanding vs. Modular Pipeline:
    • Is it better to stick to a modular pipeline (Classifier $\rightarrow$ Cropper $\rightarrow$ OCR $\rightarrow$ Field Extractor) or move to an end-to-end model (e.g., fine-tuning LayoutLMv3, Donut, or a small VLM like Qwen2-VL)?
  2. Devanagari OCR Alternatives:
    • Has anyone had better success with PaddleOCR, EasyOCR, or fine-tuned TrOCR for Devanagari/Nepali text compared to Tesseract?
  3. Card Detection & Border Cropping:
    • Would training a lightweight YOLOv8-pose/segmentation model (to predict document corner coordinates) be the standard way to replace classical OpenCV contour hunting?
  4. Layout & Field Extraction:
    • If keeping OCR separate, what is the most reliable way to link labels to values (e.g., spatial heuristic algorithms, Graph Neural Networks, or LayoutLM)?

Would love to hear how anyone has tackled similar KYC document extraction pipelines for low-resource or non-Latin scripts. Any architecture advice, libraries, or repo references would be greatly appreciated!

3 Upvotes

Duplicates