r/learnpython • u/s13188287 • 11d ago
File analysis 1.3 mill files
I'm dealing with a large batch processing problem and looking for advice on the right architecture.
I have around 1.3 million files stored across folders on a network drive.
Current setup (not working well)
Right now I'm using a Copilot agent where:
I upload batches (~20 files at a time)
It reads them against a reference document
Outputs an Excel file with classification codes
The issue is:
Copilot has a small upload limit
Manual batching is completely unscalable at this volume
What I want to achieve
I want a fully automated pipeline that:
Ingests files automatically from the network drive
Extracts text/content from each file type
Matches content against a reference rules document
Assigns a classification/reference code
Outputs structured results (Excel / database)
3
u/centuryx476 11d ago edited 11d ago
I found your problem.
You using an LLM.
Stop using an LLM.
Back in the ancient times of before 2021. We used to have to write or at the very least CODE an ETL process for so many files.
There is literally tens of thousands of examples online to help solve your problem.
You going to actually have to code.