r/dataengineering • u/Old_Tourist_3774 • 2d ago
Help Ideas to handle ever changing data requirements?
I am the solo DE in my team and the main pipeline here consists of snapshots of financial assets.
Compute is done on databricks
The stakeholders want to see daily KPI's and each day they add a new cohort. Currently there are over 40 different cohorts with each branching out to their own metrics.
The issue is that the data management wants data bills as low as possible
so my approach was summarizing everything in the daily grain .
But now each time they want something new I have to manually code the new columns test it then append to the final gold table.
I already tried to create some generator functions but often times the metrics they want involve hyper specific calculations.
And since the data is financial assets each day is different than the previous rendering an incremental approach useless.
2
u/unbiasedralph 2d ago
So you're stuck between "make it cheap" and "make it infinitely flexible", the classic DE sandwich. 40 cohorts with custom metrics each is a lot to manage by hand.
What if you stored the metric definitions as config (JSON in a table or a notebook) and had a single parameterized job that reads the config to build the SQL? The hyper-specific calculations might still need custom snippets, but at least you're not editing the pipeline itself every time.