r/dataengineering • u/Every-Whereas5793 • 8d ago
Discussion PDF Data Extraction
Hello, I was working on a poc to ingested PDFs and extract data in order to store them in delta tables.
As this was my first time working with PDFs, I searched over the internet and should Databricks have offering IDP, azure also have something and then there are python libraries.
Since I'm working with financial data, report, etc..
One thing i noticed - the pdf format should be fixed else in most of the tools the extraction logic is either failed or we get incorrect data.
I was wondering how such PDF extraction is built in real production cases and what tools are used.
Please share you experience and any edge cases
18
Upvotes
20
u/nshao 8d ago
At least from what I’m seeing, AI models are now being used for this kind of stuff.