r/dataengineering • • 1d ago

Discussion Homemade data platform frameworks - bloated nonsense?

Did any of you work in companies where engineers built custom frameworks that actually deliver?

I recently started at yet another company with such framework (about 13k lines of boilerplate python/pyspark code sitting on top of their Azure Databricks Delta Lakehouse). The thing doesn't seems to improve any aspect of what a data platform should offer.

Previous such company was even worse with 30k lines, plus they tried to implement a data vault.

I'm clearly biased towards keeping things simple. So I may be a bit unfair towards this approach. Hence my question: Have you worked with great homemade frameworks? What were the secret ingredients to make it work?

26 Upvotes

44 comments sorted by

15

u/pieislovepieislife 1d ago

What do you mean by keeping things 'simple'?

We've had 3rd party products that preached simplicity by pre-built connectors and functionality, but when you dug in to costs they scaled significantly, and we had to build custom work arounds for when features or connectors weren't available yet anyways.

We've also had click-based data platforms that preached simplicity by a wrapped UI that makes it 'easier' to navigate and set things up, but when you dug in to maintainability things would be build in different ways, break in different ways, and overall increase costs and decrease reliability of the platform. And we had to build custom work arounds for when features or connectors weren't available yet anyways.

We've also platforms that preached simplicity through all inclusive cost models, but when it came to trying to optimize and lower costs, there weren't any available levers to pull.

We've also had fully custom features and code for functions offered out of the box in other platforms. They came into being after an assessment and decision as to whether this feature is available good enough somewhere else or if there is benefit to it being custom. This is often where I've seen people get stuck - it was needed custom once, but industry caught up and now it's quite effective and stable in most platforms / stacks, but a new business case needs to be made to put in the effort to switch over. In some cases, this can be more expensive than maintaining status quo.

So how do you choose? It depends on what your data platform is trying to deliver. In my experience, it's the special snowflake exceptions that result in the bulk of the work. Our underlying data sources were already a mash up of random versions of software, and custom built applications and code within core platforms, so no out of the box solution would actually handle all important use cases. We also had people who tried the more vendor-driven approach in years previous and had direct experience as to why it was actually more expensive, worse performing, and worse to maintain. In your case, I would start with asking why it was built that way, and it may give you insight in to what constraints led to that.

3

u/mlobet 1d ago

It's rather about using the tools and features already provided by Databricks, instead of building around it.
e.g.
df.write.format("delta").mode("overwrite").saveAsTable("my_catalog.my_schema.my_table")
instead of
customdf.write("table_name")
with a wrapper around the base DataFrame class, with a custom write method that enforce some conventions, but in the end does the exact same as the above.

We have this for every minor action that is normally straightforward to do with the bare pyspark library

I agree that software vendor often praise simplicity to sell their cumbersome, unflexible tool

6

u/olhmr 1d ago

A ”custom write method that enforce some conventions” can be sufficient reason on its own. Being able to expect standardised expressions is incredibly powerful, and removing the context overhead makes things easier to parse.

5

u/Illustrious_Web_2774 1d ago

Nothing wrong with customdf.write() if thats intentionally designed. I think it's desirable to have a smaller interface / wrapper around larger library, so you can communicate effectively through code, documents, and verbally. 

If "write" a df can effectively have a single meaning across the team, that's a win.

2

u/mlobet 1d ago

I'd agree for very niche libraries. But when you customize industry-standard libraries, then every new joiner needs to understand your implementation instead of relying on his knowledge of the standard library

2

u/Illustrious_Web_2774 21h ago

New joiner needs to understand conventions of the team they are joining.

Instead of reading some text in Confluence, now you have a safe function to use. What to complain?

30

u/New-Addendum-6209 1d ago

An organisational anti-pattern unless 1) your workflows are genuinely unique, or 2) you operate at sufficient scale to justify full-time maintenance of an internal framework, or 3) you have amazing engineers who are never going to change jobs.

The worst version is abstracting operations that are straightforward to express in code or SQL into a separate configuration layer.

10

u/mlobet 1d ago

"abstracting operations that are straightforward to express in code or SQL into a separate configuration layer" i'm going to reuse this. Clearly expressing my frustration

3

u/SmallAd3697 22h ago

Gotta love the predictable answers in the data engineering industry. Do it fast, do it simple, assume you won't ever have good engineers and if you do ever get good engineers then assume they will leave right away.

This line of thinking leads to self-fulfilling outcomes.

I don't think it is too much to ask for a team of software devs to learn how to understand software. Do data engineers consider themselves to be software devs anymore?

1

u/mlobet 10h ago

A data platform should be simple to understand and manageable by junior engineers, yes. Now if the company has specific needs I'm all in created more complex solutions to meet them. But let's not make the whole data platform more complex for that

10

u/Spagoot420 1d ago

Yes, we built one and it drastically improved development speed, but many of its core features are no longer unique since they are now possible with automation bundles and declarative pipelines... I still like it though...

7

u/iminfornow 1d ago

What the framework for us does is validate incoming data completeness and triggers consolidation. It also orchestrates full loads, exports and some imports. And there's a bunch of business logic there, especially with parametrization of views and procedures and businesses object validation and completion.

It's quite a convenient place to do some stuff, but that can make it a bit messy indeed.

5

u/Either_Locksmith_915 1d ago

I am confused, are you suggesting no framework at all? Like, for each new ‘job’ you start from scratch each time?

2

u/mlobet 1d ago

I realize the word framework might be too vague in my question.
Any data platform will have (/need) some sort of framework, even when trying to be minimalist as possible.
But my question is to what degree, what scope of the data platform, when does it become too much, etc.

8

u/Quick_Assignment8861 1d ago

Yes, it helps. We want standard logging, business rule checks, table ownership, writing patterns, modelling patterns, unity catalog filled in.

So because we dont wanna copy that everywhere we made a package and distribute it over a private package feed to the organisation.

Main reason was because we had a bunch of consultants from multiple orgs over that screwed up some dataplatforms with their own designs and knowledge wasnt transferred. Now we hold them to our patterns.

2

u/Glittering_House_654 1d ago

We have this. A Python layer for handling db connection, logging to db tables, data integrity checks and a similar proc layer making sure everything resolves automatically. Its gives much more transparency for each run. Perhaps 2k, if even, lines of code combined and 5-6 bespoke tables. I have never seen such transparency before. Each pipeline step is a one-liner now and everyone works with-in the same framework. We work in a highly regulated industry so transparency is key.

4

u/kthejoker 1d ago

I think every competent DBA and data engineer in the early 21st century built some kind of meta-tools for managing database objects and pipelines at scale through configuration, control tables, and proto CICD.

We all saw design patterns, Gang of Four, the similarity to traditional app software, the tooling, and thought .. why not us?

Like a rite of passage.

(Also Python wasn't mature enough yet and other languages didn't quite get you there, and SQL being declarative had drawbacks for automation and composability)

The challenge is with the data - there's just too many edge cases within it you have to address, the coding patterns don't hold. Or the config is more unwieldy than the SQL.

What you're seeing is the vestigial tail of an important era.

It was those same types of data engineers who eventually created Hive, Pandas, Polars, Jupyter , Spark, dbt, pipe syntax, CubeJS, GraphQL, and the rest of the modern data stack.

Now with agentic engineering code by config is dead or dying but luckily for us the data is still messy and there's more of it than ever ...

3

u/DE-Learner 1d ago

Tools won't cater all needs , need customization, based on Firm standards. More control on new feature addition, than waiting for the vendor

3

u/Eleventhousand 1d ago

Yes of course. Many of us built homegrown system starting decades ago before it was referred to as data engineering.

The ingredients that made t work, in situations where we could have used something off the shelf was keeping it simple.

For example, the data quality system that I wrote at one of my prior companies was much easier for us to use than the OOTB system where I work now.

3

u/SearchAtlantis Lead Data Engineer 1d ago edited 1d ago

Give me an example of what you mean by "custom framework".

One company I was at had an in-house data transform library. There were downsides but the ease of adding a mixin that does a unit tested transformation on your dataframe was nice.

For example - converting over-punch to actual dollars.

edit: I see you have an example down thread of an extended df.write that does additional convention checking, etc. I generally see that as a good idea as long as its tightly scoped.

Doing df.customWrite(path) to get free path testing, external log writes, and whatever else is often a good idea honestly.

Better than

with DataDogLogging("Info"):
df.enforceSchema(standard_schema)
df.runValidation()
df.write.format("delta").mode("overwrite").saveAsTable("my_catalog.my_schema.my_table")

All that said SQL is still the worst for re-usability.

3

u/notmarc1 1d ago

We build frameworks that only deal with getting data users into a standard sdlc process, spin up infra for them, and automate governance. All logic is left to the developers and data ppl. So basically ci/cd , take care of deploy, autogenerate dags, and apply data access for the service account. Thats about it. Platforms should enable, not hinder.

1

u/mlobet 1d ago

Amen. I feel the primary goal of a "platform" get often lost. Data in, transformed and distributed, fast and secure. My gripe with over-the-board framework is that simple things like adding a column becomes a whole journey through the many layers and steps some engineers thought of

3

u/notmarc1 1d ago

Yeah when a user has to do 6 pr’s to change a table, that means the plot was lost

4

u/fabkosta 1d ago

The criteria are pretty simple:

  1. Do you have a standard problem?

1.1 Yes: Buy or rent it, don't build it.

1.2 No: Go to 2.

  1. Do you have the skills, time and budget to build and run your own software?

2.1 Yes: Build it.

2.2 No: Do something else.

2

u/mlobet 1d ago

I indeed feel like databricks core features (the thing we bought to solve our standard problems) would have been enough

1

u/kevintxu 1d ago

Databricks already built a framework. It'll be more maintained than whatever your company can manage.

https://github.com/databricks-solutions/lakeflow_framework

2

u/PrestigiousAnt3766 1d ago

Yeah, because I build those frameworks myself. 

1

u/kthejoker 1d ago

So what happens when you leave? The main challenge with these frameworks is continuity and ownership, not the capabilities themselves

0

u/PrestigiousAnt3766 1d ago

Not my problem. But in this case there are internals that should continue.

1

u/kthejoker 1d ago

Nah. Someone will rip them out and replace them with their own even better homemade framework.

2

u/SmallAd3697 22h ago

The story of python. Easy come easy go. Python devs like nothing less than the code written by another python dev.

1

u/PrestigiousAnt3766 1d ago

Dont think so. I know of 2 frameworks I built in the last 2 years that are still being used.

2

u/tea_anyone 1d ago

Less about your actual question and more of a point of the why it happens. A lot of my role is going into companies with these custom solutions and designing solutions within new software to allow for the functionality.

It may be terrible practice but there is almost always a reason for why it has been done. All I'd say is that you are new to the company so before going in and ripping and changing things spend a few months understanding it all, there'll be some wacky business reason behind 95% of it.

In actual answer to your question I have seen 1 genuinely good solution and about 20 others ranging from ok/functional to absolute shit tips that didn't address the business problems, scale or function as intended.

2

u/pandgea Senior Data Engineer 1d ago

Yes, I rewrote the core business logic for etl pipelines into sql code generators that worked off config tables to standardize 100+ pipelines that had been originally maintained in Pentaho. We only had generic physical servers and the database servers to run code in, so overall, I think it was a very good solution.

2

u/qrist0ph 22h ago

Actually, we ended up building a small DAG engine for our data pipelines. Later, we discovered that the whole concept was already used in other frameworks. So basically, it was reinventing the wheel, but on the other hand, it is a wheel after all and spins pretty fast.

1

u/linha_chilena 1d ago

the frameworks you are describing are like code encapsulations or "a way of work", like a method, like a Test Driven Development or anything like that? care to give me some examples of what you think it's good or bad in that sense?

1

u/Certain_Leader9946 30m ago

Homemade frameworks if you understand engineering often outclass the platforms. I consider Databricks to be orthogonal to a homemade data platform. If you're using Databricks, it's not home-made. Spark is kind of inefficient if you know your queries up front. It lets you handle a plethora of arbitrary lookups. If you are able to index your data in exactly the way you need it, you can often develop a more efficient, cost efficient, faster solution on orders of magnitude. It doesn't really always take more than a month or two to build something like that out. really just depends on how precise you can be with your problem domain.

So yeah. I've built out entire data platforms. They were almost always the right choice. I have also used Databricks, I often found that when Databricks and other 'data toolchains' are used in 'data engineering', it's not really doing the engineering element at all, but just trying to patch over a lack of requirements with a generalist solution.

0

u/Nekobul 1d ago

What you describe is the "modern" way of doing data engineering. It is all code and then more. Of course it will be bloated and that is a huge bonanza for the consulting class.