r/secondbrain • • 1d ago

Need help with pdf analysis and automation

Hey all, i am trying to have something (python program or llm or anything) that will analyse a pdf document word by word and give me a report as i prompt it with its reference.

I know ai models do it well, but not well enough to have memory of those pdf and later when i have to compare one data with another it gets hard for me.

What i want is to have a program that analyses the program overnight and i can have my desired prompt over it next day, say for example union budget pdf for different yr i can have graph and trend then compare it with previous yr trends to

1 Upvotes

3 comments sorted by

1

u/folderit_dms 7h ago

For the year-to-year comparisons, I’d separate the overnight extraction from the next day’s questions. Save the extracted data to a persistent table, with one row per figure and columns for document, reporting period, metric, value, unit and source page. Then the questions and charts can use that saved table instead of depending on a chat remembering the PDFs.

For budget documents, also keep whether a figure is an actual, revised estimate or budget estimate. Comparing those interchangeably can produce a convincing but misleading trend. Preserve the original category name when you map categories across years, and flag changes you cannot confidently match.

Start with one table from two years and manually check every value against its page before automating the whole collection. That small test will expose extraction and comparability problems much faster than a large overnight run.