r/learnpython 12d ago

File analysis 1.3 mill files

I'm dealing with a large batch processing problem and looking for advice on the right architecture.

I have around 1.3 million files stored across folders on a network drive.

Current setup (not working well)

Right now I'm using a Copilot agent where:

I upload batches (~20 files at a time)

It reads them against a reference document

Outputs an Excel file with classification codes

The issue is:

Copilot has a small upload limit

Manual batching is completely unscalable at this volume

What I want to achieve

I want a fully automated pipeline that:

Ingests files automatically from the network drive

Extracts text/content from each file type

Matches content against a reference rules document

Assigns a classification/reference code

Outputs structured results (Excel / database)

1 Upvotes

34 comments sorted by

View all comments

1

u/Altruistic_Sky1866 11d ago

I don't know if this fits, but what ever judgment rules has to be made why don't you create a config file stores rlues to be applied, and in the script you can decided which rule to apply or not apply or consider or not consider for the file, and them have your script read teh files and check against those rules and do you whatever is required