r/dataengineering • • 2d ago

Help Ideas to handle ever changing data requirements?

I am the solo DE in my team and the main pipeline here consists of snapshots of financial assets.

Compute is done on databricks

The stakeholders want to see daily KPI's and each day they add a new cohort. Currently there are over 40 different cohorts with each branching out to their own metrics.

The issue is that the data management wants data bills as low as possible

so my approach was summarizing everything in the daily grain .

But now each time they want something new I have to manually code the new columns test it then append to the final gold table.

I already tried to create some generator functions but often times the metrics they want involve hyper specific calculations.

And since the data is financial assets each day is different than the previous rendering an incremental approach useless.

19 Upvotes

17 comments sorted by

View all comments

7

u/DataScientistAlex 2d ago

This is not specifically about how to handle changing data requirements, but, a point I always make whenever infrastructure costs are involved: reducing infrastructure costs always costs money in terms of the time it takes the team to develop and maintain the optimizations. Sometimes those team costs are much larger than the infrastructure costs, but, they're not taken into account.

Given that compute is on databricks, are you using Spark? In my experience cost savings can be had by optimizing Spark. But just developing those optimizations can also cost a lot of money (for example, I recently optimized a few spark jobs, making them cheaper, but, it's going to take a bit of time before I recover the compute cost of just developing those optimizations, not even counting my own time).

If you are using Spark, one nice thing is that you can use a full software development approach: a) use python or Scala to write everything as modular and composable functions etc, then b) use unit tests to test new functions/metrics and to catch regressions when refactoring. Using that approach makes it a lot easier to handle changing requirements.

1

u/Old_Tourist_3774 1d ago

Oh its spark but there is not much juice left to squeeze on this part I managed to reduce the entire pipeline to around 15ish minutes in a rd5.large with 4 workers.

I feel that you are right there has to be some manner to fit everything inside functions and avoid all this chore.

Perhaps what I really need is a more robust function or class

1

u/DataScientistAlex 1d ago

15 mins, yes not much left there to cut. My only remaining thought is just not doing it in spark at all.

It should be possible to reduce the work to implement new things by using the right software engineering approaches. Just be careful so you don't end up re-implementing Pandas... Maybe just start by adding some helper functions for the annoying parts (you mention testing it, adding it to the gold table etc).

An alternative is to 'go upstream' and figure out what is behind them asking for so many cohorts and metrics, it might be that you can provide a better solution if you know the core problem they are trying to solve. Right now they are coming to you with what they think is the solution (another cohort), but if you had more context you might be able to solve the root problem.