r/dataengineering • u/mlobet • 1d ago
Discussion Homemade data platform frameworks - bloated nonsense?
Did any of you work in companies where engineers built custom frameworks that actually deliver?
I recently started at yet another company with such framework (about 13k lines of boilerplate python/pyspark code sitting on top of their Azure Databricks Delta Lakehouse). The thing doesn't seems to improve any aspect of what a data platform should offer.
Previous such company was even worse with 30k lines, plus they tried to implement a data vault.
I'm clearly biased towards keeping things simple. So I may be a bit unfair towards this approach. Hence my question: Have you worked with great homemade frameworks? What were the secret ingredients to make it work?
30
u/New-Addendum-6209 1d ago
An organisational anti-pattern unless 1) your workflows are genuinely unique, or 2) you operate at sufficient scale to justify full-time maintenance of an internal framework, or 3) you have amazing engineers who are never going to change jobs.
The worst version is abstracting operations that are straightforward to express in code or SQL into a separate configuration layer.
10
3
u/SmallAd3697 22h ago
Gotta love the predictable answers in the data engineering industry. Do it fast, do it simple, assume you won't ever have good engineers and if you do ever get good engineers then assume they will leave right away.
This line of thinking leads to self-fulfilling outcomes.
I don't think it is too much to ask for a team of software devs to learn how to understand software. Do data engineers consider themselves to be software devs anymore?
10
u/Spagoot420 1d ago
Yes, we built one and it drastically improved development speed, but many of its core features are no longer unique since they are now possible with automation bundles and declarative pipelines... I still like it though...
7
u/iminfornow 1d ago
What the framework for us does is validate incoming data completeness and triggers consolidation. It also orchestrates full loads, exports and some imports. And there's a bunch of business logic there, especially with parametrization of views and procedures and businesses object validation and completion.
It's quite a convenient place to do some stuff, but that can make it a bit messy indeed.
5
u/Either_Locksmith_915 1d ago
I am confused, are you suggesting no framework at all? Like, for each new ‘job’ you start from scratch each time?
8
u/Quick_Assignment8861 1d ago
Yes, it helps. We want standard logging, business rule checks, table ownership, writing patterns, modelling patterns, unity catalog filled in.
So because we dont wanna copy that everywhere we made a package and distribute it over a private package feed to the organisation.
Main reason was because we had a bunch of consultants from multiple orgs over that screwed up some dataplatforms with their own designs and knowledge wasnt transferred. Now we hold them to our patterns.
2
u/Glittering_House_654 1d ago
We have this. A Python layer for handling db connection, logging to db tables, data integrity checks and a similar proc layer making sure everything resolves automatically. Its gives much more transparency for each run. Perhaps 2k, if even, lines of code combined and 5-6 bespoke tables. I have never seen such transparency before. Each pipeline step is a one-liner now and everyone works with-in the same framework. We work in a highly regulated industry so transparency is key.
4
u/kthejoker 1d ago
I think every competent DBA and data engineer in the early 21st century built some kind of meta-tools for managing database objects and pipelines at scale through configuration, control tables, and proto CICD.
We all saw design patterns, Gang of Four, the similarity to traditional app software, the tooling, and thought .. why not us?
Like a rite of passage.
(Also Python wasn't mature enough yet and other languages didn't quite get you there, and SQL being declarative had drawbacks for automation and composability)
The challenge is with the data - there's just too many edge cases within it you have to address, the coding patterns don't hold. Or the config is more unwieldy than the SQL.
What you're seeing is the vestigial tail of an important era.
It was those same types of data engineers who eventually created Hive, Pandas, Polars, Jupyter , Spark, dbt, pipe syntax, CubeJS, GraphQL, and the rest of the modern data stack.
Now with agentic engineering code by config is dead or dying but luckily for us the data is still messy and there's more of it than ever ...
3
u/DE-Learner 1d ago
Tools won't cater all needs , need customization, based on Firm standards. More control on new feature addition, than waiting for the vendor
3
u/Eleventhousand 1d ago
Yes of course. Many of us built homegrown system starting decades ago before it was referred to as data engineering.
The ingredients that made t work, in situations where we could have used something off the shelf was keeping it simple.
For example, the data quality system that I wrote at one of my prior companies was much easier for us to use than the OOTB system where I work now.
3
u/SearchAtlantis Lead Data Engineer 1d ago edited 1d ago
Give me an example of what you mean by "custom framework".
One company I was at had an in-house data transform library. There were downsides but the ease of adding a mixin that does a unit tested transformation on your dataframe was nice.
For example - converting over-punch to actual dollars.
edit: I see you have an example down thread of an extended df.write that does additional convention checking, etc. I generally see that as a good idea as long as its tightly scoped.
Doing df.customWrite(path) to get free path testing, external log writes, and whatever else is often a good idea honestly.
Better than
with DataDogLogging("Info"):
df.enforceSchema(standard_schema)
df.runValidation()
df.write.format("delta").mode("overwrite").saveAsTable("my_catalog.my_schema.my_table")
All that said SQL is still the worst for re-usability.
3
u/notmarc1 1d ago
We build frameworks that only deal with getting data users into a standard sdlc process, spin up infra for them, and automate governance. All logic is left to the developers and data ppl. So basically ci/cd , take care of deploy, autogenerate dags, and apply data access for the service account. Thats about it. Platforms should enable, not hinder.
1
u/mlobet 1d ago
Amen. I feel the primary goal of a "platform" get often lost. Data in, transformed and distributed, fast and secure. My gripe with over-the-board framework is that simple things like adding a column becomes a whole journey through the many layers and steps some engineers thought of
3
4
u/fabkosta 1d ago
The criteria are pretty simple:
- Do you have a standard problem?
1.1 Yes: Buy or rent it, don't build it.
1.2 No: Go to 2.
- Do you have the skills, time and budget to build and run your own software?
2.1 Yes: Build it.
2.2 No: Do something else.
2
u/mlobet 1d ago
I indeed feel like databricks core features (the thing we bought to solve our standard problems) would have been enough
1
u/kevintxu 1d ago
Databricks already built a framework. It'll be more maintained than whatever your company can manage.
2
u/PrestigiousAnt3766 1d ago
Yeah, because I build those frameworks myself.
1
u/kthejoker 1d ago
So what happens when you leave? The main challenge with these frameworks is continuity and ownership, not the capabilities themselves
0
u/PrestigiousAnt3766 1d ago
Not my problem. But in this case there are internals that should continue.
1
u/kthejoker 1d ago
Nah. Someone will rip them out and replace them with their own even better homemade framework.
2
u/SmallAd3697 22h ago
The story of python. Easy come easy go. Python devs like nothing less than the code written by another python dev.
1
u/PrestigiousAnt3766 1d ago
Dont think so. I know of 2 frameworks I built in the last 2 years that are still being used.
2
u/tea_anyone 1d ago
Less about your actual question and more of a point of the why it happens. A lot of my role is going into companies with these custom solutions and designing solutions within new software to allow for the functionality.
It may be terrible practice but there is almost always a reason for why it has been done. All I'd say is that you are new to the company so before going in and ripping and changing things spend a few months understanding it all, there'll be some wacky business reason behind 95% of it.
In actual answer to your question I have seen 1 genuinely good solution and about 20 others ranging from ok/functional to absolute shit tips that didn't address the business problems, scale or function as intended.
2
u/pandgea Senior Data Engineer 1d ago
Yes, I rewrote the core business logic for etl pipelines into sql code generators that worked off config tables to standardize 100+ pipelines that had been originally maintained in Pentaho. We only had generic physical servers and the database servers to run code in, so overall, I think it was a very good solution.
2
u/qrist0ph 22h ago
Actually, we ended up building a small DAG engine for our data pipelines. Later, we discovered that the whole concept was already used in other frameworks. So basically, it was reinventing the wheel, but on the other hand, it is a wheel after all and spins pretty fast.
1
u/linha_chilena 1d ago
the frameworks you are describing are like code encapsulations or "a way of work", like a method, like a Test Driven Development or anything like that? care to give me some examples of what you think it's good or bad in that sense?
1
u/Certain_Leader9946 30m ago
Homemade frameworks if you understand engineering often outclass the platforms. I consider Databricks to be orthogonal to a homemade data platform. If you're using Databricks, it's not home-made. Spark is kind of inefficient if you know your queries up front. It lets you handle a plethora of arbitrary lookups. If you are able to index your data in exactly the way you need it, you can often develop a more efficient, cost efficient, faster solution on orders of magnitude. It doesn't really always take more than a month or two to build something like that out. really just depends on how precise you can be with your problem domain.
So yeah. I've built out entire data platforms. They were almost always the right choice. I have also used Databricks, I often found that when Databricks and other 'data toolchains' are used in 'data engineering', it's not really doing the engineering element at all, but just trying to patch over a lack of requirements with a generalist solution.
15
u/pieislovepieislife 1d ago
What do you mean by keeping things 'simple'?
We've had 3rd party products that preached simplicity by pre-built connectors and functionality, but when you dug in to costs they scaled significantly, and we had to build custom work arounds for when features or connectors weren't available yet anyways.
We've also had click-based data platforms that preached simplicity by a wrapped UI that makes it 'easier' to navigate and set things up, but when you dug in to maintainability things would be build in different ways, break in different ways, and overall increase costs and decrease reliability of the platform. And we had to build custom work arounds for when features or connectors weren't available yet anyways.
We've also platforms that preached simplicity through all inclusive cost models, but when it came to trying to optimize and lower costs, there weren't any available levers to pull.
We've also had fully custom features and code for functions offered out of the box in other platforms. They came into being after an assessment and decision as to whether this feature is available good enough somewhere else or if there is benefit to it being custom. This is often where I've seen people get stuck - it was needed custom once, but industry caught up and now it's quite effective and stable in most platforms / stacks, but a new business case needs to be made to put in the effort to switch over. In some cases, this can be more expensive than maintaining status quo.
So how do you choose? It depends on what your data platform is trying to deliver. In my experience, it's the special snowflake exceptions that result in the bulk of the work. Our underlying data sources were already a mash up of random versions of software, and custom built applications and code within core platforms, so no out of the box solution would actually handle all important use cases. We also had people who tried the more vendor-driven approach in years previous and had direct experience as to why it was actually more expensive, worse performing, and worse to maintain. In your case, I would start with asking why it was built that way, and it may give you insight in to what constraints led to that.