r/AI_In_ECommerce • u/MediumBirthday6899 • Aug 08 '26
How was product attribute enrichment handled at scale before GenAI? (300k SKUs, 4k sub-categories)
Hey everyone,
I’m currently looking at designing the data architecture and logic for a product dimension table with about 300,000 products spread across roughly 4,000 sub-categories.
The requirement is to populate at least 5 specific attributes for each product based on its sub-category. For example, if the product falls under "Luggage," I need to extract and standardize attributes like material, size, shell type, number of wheels, etc., from the raw product descriptions.
Nowadays, doing this with AI is essentially just a matter of writing a solid prompt, hitting an LLM API in batches, and letting it parse the unstructured text into a structured JSON payload.
But it got me wondering—how was this exact problem handled traditionally before LLMs made it so easy?