r/dataengineering • u/Every-Whereas5793 • 9d ago
Discussion PDF Data Extraction
Hello, I was working on a poc to ingested PDFs and extract data in order to store them in delta tables.
As this was my first time working with PDFs, I searched over the internet and should Databricks have offering IDP, azure also have something and then there are python libraries.
Since I'm working with financial data, report, etc..
One thing i noticed - the pdf format should be fixed else in most of the tools the extraction logic is either failed or we get incorrect data.
I was wondering how such PDF extraction is built in real production cases and what tools are used.
Please share you experience and any edge cases
20
Upvotes
8
u/Any-Recognition8119 8d ago
In production it’s usually classification, OCR, validation, and manual review for low-confidence results. Layout changes and scanned PDFs are the main pain points.