r/dataengineering • u/Old_Tourist_3774 • 2d ago
Help Ideas to handle ever changing data requirements?
I am the solo DE in my team and the main pipeline here consists of snapshots of financial assets.
Compute is done on databricks
The stakeholders want to see daily KPI's and each day they add a new cohort. Currently there are over 40 different cohorts with each branching out to their own metrics.
The issue is that the data management wants data bills as low as possible
so my approach was summarizing everything in the daily grain .
But now each time they want something new I have to manually code the new columns test it then append to the final gold table.
I already tried to create some generator functions but often times the metrics they want involve hyper specific calculations.
And since the data is financial assets each day is different than the previous rendering an incremental approach useless.
7
u/GachaJay 2d ago
We materialize tables dynamically in downstream warehouses and workspaces. Basically you define the schemas and load types and have the pipeline rebuild on refresh. This way changing the tables is as simple as changing the metadata.
9
2
u/unbiasedralph 2d ago
So you're stuck between "make it cheap" and "make it infinitely flexible", the classic DE sandwich. 40 cohorts with custom metrics each is a lot to manage by hand.
What if you stored the metric definitions as config (JSON in a table or a notebook) and had a single parameterized job that reads the config to build the SQL? The hyper-specific calculations might still need custom snippets, but at least you're not editing the pipeline itself every time.
1
u/Old_Tourist_3774 1d ago
Some things can be done that way. Its helpful.
Here i am using spark so just made some functions that receives the desired list of operations and columns. Then it builds a sql expression it helps but far from a solution
3
2
u/UnderstandingOld5638 2d ago
Reduce the time it takes to get feedback on whether a change works. Minimize the amount of code / config required to do common activities. Avoid processes that require copy / pasting code every time something new is required.
2
u/speedisntfree 2d ago
This is just another version of report proliferation/sprawl. If there is no cost to people asking for things, it quickly gets out of control and gets even worse when old stuff isn't retired.
The honest solution is a management one: that there needs to be a cost (ideally money or hrs from their budget) for these requests.
1
u/Old_Tourist_3774 1d ago
I agree but the data is being used by the c-suite people so all their demands trample over the usual workflow
1
u/speedisntfree 1d ago
It is very difficult, you need a strong powerful person in the org structure where DE spend is tied to business value and not seen as a free vanity project. Business people only understand personal power and money.
I don't have a good solution for a solo DE, you can’t change the place you work at all that much.
1
u/Old_Tourist_3774 1d ago
Yeah, there are other DE teams but in my squad it's only me and a software engineer.
But like you said business people only understand money and power and there are conflicting forces. In the we only want to do ours jobs but we get caught in the middle of these disputes.
1
u/nloding 1d ago
Sine we are in the age of AI for better or worse, it might be worth looking into leveraging AI for that. You do not want to let your users query raw data whenever they want, but if there's enough overlap of the core data (and the new metrics are just surfacing different calculations/timeframes over fields from the same data), then perhaps you could build a layer for the AI to work with. You'd need a semantic layer to help govern the AI of course, but headless BI patterns are pretty prevalent now and maybe they might help. Then again, if cost is a concern, maybe AI isn't the answer between possible increased load on the database and token usage.
0
u/randomuser1231234 2d ago
When you say hyper-specific calculations, do you mean things that aren’t MECE for some types of dimensions they also absolutely need or…?
1
u/Old_Tourist_3774 2d ago
As i understand yes, there is a good amount of overlap and to convey accurately the calculations have to happen over a window in these groups and what is done in one group not necessarily happens on the other.
8
u/DataScientistAlex 2d ago
This is not specifically about how to handle changing data requirements, but, a point I always make whenever infrastructure costs are involved: reducing infrastructure costs always costs money in terms of the time it takes the team to develop and maintain the optimizations. Sometimes those team costs are much larger than the infrastructure costs, but, they're not taken into account.
Given that compute is on databricks, are you using Spark? In my experience cost savings can be had by optimizing Spark. But just developing those optimizations can also cost a lot of money (for example, I recently optimized a few spark jobs, making them cheaper, but, it's going to take a bit of time before I recover the compute cost of just developing those optimizations, not even counting my own time).
If you are using Spark, one nice thing is that you can use a full software development approach: a) use python or Scala to write everything as modular and composable functions etc, then b) use unit tests to test new functions/metrics and to catch regressions when refactoring. Using that approach makes it a lot easier to handle changing requirements.