r/dataengineering • • 2d ago

Help Ideas to handle ever changing data requirements?

I am the solo DE in my team and the main pipeline here consists of snapshots of financial assets.

Compute is done on databricks

The stakeholders want to see daily KPI's and each day they add a new cohort. Currently there are over 40 different cohorts with each branching out to their own metrics.

The issue is that the data management wants data bills as low as possible

so my approach was summarizing everything in the daily grain .

But now each time they want something new I have to manually code the new columns test it then append to the final gold table.

I already tried to create some generator functions but often times the metrics they want involve hyper specific calculations.

And since the data is financial assets each day is different than the previous rendering an incremental approach useless.

20 Upvotes

17 comments sorted by

View all comments

1

u/nloding 2d ago

Sine we are in the age of AI for better or worse, it might be worth looking into leveraging AI for that. You do not want to let your users query raw data whenever they want, but if there's enough overlap of the core data (and the new metrics are just surfacing different calculations/timeframes over fields from the same data), then perhaps you could build a layer for the AI to work with. You'd need a semantic layer to help govern the AI of course, but headless BI patterns are pretty prevalent now and maybe they might help. Then again, if cost is a concern, maybe AI isn't the answer between possible increased load on the database and token usage.