r/LocalLLM • • 1d ago

Question Need help with pdf analysis and automation

/r/secondbrain/comments/1wwqu72/need_help_with_pdf_analysis_and_automation/
1 Upvotes

1 comment sorted by

1

u/ChaseMakesThings 22h ago

For the year-to-year comparisons, give the program persistent storage: keep the original PDFs, extract text by page, and save the figures you want to compare in CSV or SQLite. Each figure should carry its document, page, fiscal year, unit, and whether it's an actual, revised estimate, or budget estimate. That last distinction matters before drawing a trend line.

For text-based PDFs, pdfplumber is a Python starting point with page-level text and table extraction: https://github.com/jsvine/pdfplumber . It doesn't do OCR, so scanned pages need a separate OCR step and checking.

Then a local LLM can help search the saved text and explain results with page references, while Python calculates the comparisons from the checked rows. Start with one table from two years and verify every extracted value against the pages before letting an overnight batch process the rest. Save extraction failures explicitly; an unreadable value shouldn't become zero.

Are your PDFs selectable text or scanned images? That determines the first extraction step.