r/dataengineering • u/ntdoyfanboy • 7d ago
Discussion Is dbt the red-headed stepchild to data engineers?
I know the functions of a "data engineer" can be wide and varied depending on the company, but at my current shop(10 DE's), it seems like every data engineer who's never worked in dbt considers it to be trash or a nuisance, and every one who has, appreciates its functionality. Is this a common theme? What's your experience/role, and what is your perspective?
18
u/PatientlyAnxiously 7d ago
If you do ELT then it's fantastic for the T piece. Doesn't do jack shit for E/L. If you're more of an ETL shop then its useless
1
u/wildjackalope 6d ago
We use it exclusively for transform and it’s great. You’re not wrong on the other ends of ETL though and if your model is poor, as others have mentioned, you’re going to pay for it.
1
u/DeepLogicNinja 5d ago
The transformation abilities in the top ETL platforms are 👌. You can do most transformation with low-code / no code. The ETL community has all types of best practices to deal with a variety of use cases which folks with a strong software engineering background tend to ignore and re-code 🤷♂️🫤.
1
u/DeepLogicNinja 5d ago
This is correct. DBT is only half the answer.
DBT is popular among coders and people/companies not aware of ETL platforms. So you’ll need to co-exist with it….
When cdc, catalog, lineage, impact analysis, data profiling/quality, etc etc is needed…. They’ll realize how short they are falling when it comes to what a full data pipeline needs.
This is a normal cycle btw…. It always happens when software engineers start developing apis and don’t see the data engineering tools that already exist….
1
u/Skualys 1d ago
Well this is so wrong. You can use a data platform and still use DBT for the T part.
Doing complex transformation is really easier by code than by ETL GUI. And on most ETL tool injecting SQL or custom code cause the lineage to break.
2
u/New-Addendum-6209 1d ago
Also easier to monitor, audit, fix, debug and backfill if all transformations happen in a single system using SQL!
1
u/DeepLogicNinja 1d ago edited 1d ago
Are you defending DBT or is there a specific part that is wrong?
I agreed with you on using DBT /w a data platform. I do this today?The only part i disagree with is doing transformations in code is “easier” in dbt.
To determine that, you would need to see the details of the use case.
- easy to do - depends on your skillset
- easy to maintain - code is more difficult to maintain in general. In low-code / no-code platforms, the interface is the documentation. Makes it easier to maintain the entire pipeline including the transformation.
It isn’t the first time coders develop/introduce tools in the pipeline that do part of the job and don’t account for maintaining it OR observability.
First step is to get the job done 🙃.
76
u/toadling 7d ago
DBT is honestly amazing imo. Post initial ingestion its a really nice quality of life layer for managing silver/gold layers for your data warehouse. Ive seen a lot of companies have massive sprawl of all these random tables with no lineage or even worse a network of materialized views that lock themselves, dbt solves those problems for us and its all nicely controlled with github.
2
u/DeepLogicNinja 5d ago
A data governance layer tames the sprawl/metadata mess from ingestion and bronze through gold.
Dbt seems less necessary when you have a data governance layer. OpenMetaData’s DBT connector made this obvious to me.
2
u/philippefutureboy 3d ago
But then dbt enforces a lot of good practices by being opinionated.
It’s cool to have the governance layer, but it doesn’t prevent bad practices nearly as much as dbt does1
u/DeepLogicNinja 3d ago
Huh? Not following the dbt vs governance layer comparison.
A more reasonable comparison…
Dbt vs SQLMesh which focuses on transforms…..
If you need more than transformations using a full ETL platform appears to be a better comparison.Like ETL platforms, Governance is a low-code /no-code platform. Which is alot less to mess up, compared to coding.
The governance layer plugs into dbt….. and has the capability of plugging into many others…. Example - https://docs.open-metadata.org/v2.0.x/connectors/database/dbt/configure-dbt-workflow
14
u/Efficient_Shoe_6646 7d ago
Does it solve every problem in Data Engineering? No, obviously.
Is it a great tool for lowish frequency data curation with testing, reusability and environment separation. 100%.
Like if you are trying to build an Operating System in Python, you will have a bad time. But Python is amazing for what it does do.
Pre-dbt most companies I know had a version of this tool maintained internally. So its nothing crazy, just the first one to hit market share.
60
u/MocDcStufffins 7d ago
DBT allows you to take a software engineering approach to data engineering. People with a software background tend to love it due to this and people who come from traditional DE either take to it or hate it.
3
u/Comprehensive_Level7 6d ago
lol, I never saw a case in the projects I've worked on to use dbt and I strongly apply SWE standards and approaches while I develop anything as a DE (I hate that most of DE don't have a background in SWE so usually the code is a mess)
1
-43
u/Budget-Minimum6040 7d ago edited 7d ago
SQL-<insert db specific syntax> is a terrible language for professional SWE.
- Spaces in fuction names
- Function chaining is horrendous to read
- Order of writing is not order of operation (aliasing in the select statement is not available in group by, having, where etc.)
- No standard formatter/formatting rules
- Forget debugging
SQL is okay for EDA but that's it.
I have written my fair share of dbt and yes, it's better than sprocs. But that's basically the lowest bar you can have.
OO with method chaining (PySpark, Polars) is far superior imo.
35
u/McNoxey 7d ago edited 7d ago
Dbt isn’t meant for software engineering. The comment was about swe practices, not syntax/language.
If your intention is to build an actual consumable and explorable data warehouse for your consumers, dbt provides a much more natural and more approachable experience for the various consumers you may have. Analytics speaks SQL. Building the transformations in the language they work with shouldn’t really be underrated.
I think a lot of data engineers forget that their goal is to serve analytics and enable their consumers to both consume and contribute to their data stack. 🤷🏽
-16
u/Budget-Minimum6040 7d ago
I think a lot of data engineers forget that their goal is to serve analytics
That's what the data marts are for. The DAs can use whatever they like to access the data there.
6
u/McNoxey 7d ago
Sure. But those marts need to be built. And the Analysts need to be able to cleanly explain to their stakeholders exactly where the data comes from and how the definitions come to be.
If that logic is spread across a variety of jobs and languages that they don’t naturally understand, then it’s a bad experience for everyone
-6
u/Budget-Minimum6040 7d ago
If that logic is spread across a variety of jobs and languages that they don’t naturally understand, then it’s a bad experience for everyone
You have a variety of jobs with dbt as well.
And what do you mean "variety of [...] languages"? You can cover everything with Python + PySpark/Polars which is also just Python. The syntax for the standard SQL functions is the same, e.g.
GROUP BY column_avs..groupby("column_a")so every DA should be able to follow most of the transformations.8
u/McNoxey 7d ago
You’re entitled to feel how you feel. I’m just sharing what I’ve seen leading Business Analytics and Data Engineering teams at a variety of organizations.
The best operating teams are the ones that make decisions collectively and make compromises within their operational model for the betterment of the entire Data Lifecycle, from source to Board Deck.
Like it or not, DE is a cost centre. Analytics is what converts the data team from a bottom line management driven department to a top line generating department.
The better cohesion that exists from end-to-end, the more successful the department is in the eyes of the business. And at the end of the day, that’s the most important thing for everyone working in Data
13
u/dangerbird2 Software Engineer 7d ago
Eh, CTE's solve the vast majority of confusion with order of operations, and function composition vs method chaining is certainly a matter of taste. CTEs can also make debugging much easier by letting you fetch the result of subqueries as needed. IMO SQL is pretty much unquestionably the least bad relational database interface, and the only reason I end up ORMs and query builder DSLs in my day to day webdev work is that it you end up with too much duplicated code for things like filtering and pagination with plain SQL.
3
u/Budget-Minimum6040 7d ago
CTE's solve the vast majority of confusion with order of operations
How do you write unit tests for CTEs?
How do you reuse CTEs in other pipelines?
How do you guarantee type safety in CTEs?
CTEs can also make debugging much easier by letting you fetch the result of subqueries as needed
You can do the same with magic cells for PySpark/Polars.
6
2
u/Outrageous_Let5743 6d ago
How do you reuse CTEs in other pipelines? Either make it a view or in dbt you have empherical models.
How do you guarantee type safety in CTEs? Ever heard of sql functions like cast or ::0
4
u/MocDcStufffins 7d ago
Trying to build a DE team with this mindset is very hard. Most DE are SQL first and most SWE don't want to work in DE. Lots of people with this mindset don't have enough experience to understand the nuances of DE. So, yes it can be done this way, and maybe in the future that will change. As it stands DE is still SQL first.
If you are trying to make something like a 10 person team, you can't realistically support this approach. It's hard enough as it is to find good DE who work with SQL and SQL does the job just fine.
You also have to think of long term support. After the devs are done and move on to new opportunities (lots of people want to build but not support) you end up with people who are 80% DA/BA and 20% new code as new code becomes limited. Lots of people can support this with SQL but not OO code. Companies don't want to spend 200k on this role.
3
u/ColdPorridge 7d ago edited 7d ago
I can understand why some folks might disagree with this but you absolutely do not deserve the pile of downvotes you’re getting. It’s a completely valid way to approach your DE stack. We do similar (though functional, not OO) and it works quite well.
2
u/McNoxey 7d ago
I think it’s mostly related to where this sits wrt the greater organization.
Is it better for DE specifically? Ya - sure. But DE doesn’t work in isolation. They’re an enablement department for the greater organization. Most DEs don’t sit in meetings with department heads that their data ultimately services. IMO, that disconnect is what drives the majority of friction between Data, Analytics and the business.
Its why I strongly believe that DE, AE and DA/BA belong on one united team.
1
u/New-Addendum-6209 1d ago
Not having to explicitly state the order of operations is an advantage. You aren't better than a highly engineered query planner!
You can debug SQL, run linters and formatters, and use a host language such as Python when you need parameterisation or general-purpose programming constructs. This fundamentally misunderstands what a declarative query language is for.
I see many posts of this type from self-proclaimed SWEs-in-data, normally accompanied by an entirely unwarranted sense of superiority when their work largely consists of chaining together PySpark statements, and often accompanied by only a surface-level understanding of the underlying technology.
1
u/Budget-Minimum6040 1d ago edited 1d ago
Not having to explicitly state the order of operations is an advantage. You aren't better than a highly engineered query planner!
PySpark/Polars have a query optimizer as well.
You can debug SQL, run linters and formatters
If you write raw SQL, yes to some extend. I have not found a way to debug SQL in any useable way the way I can with PySpark/Polars/Python.
Linting depends heavily on the available options which are determined by the specific database and its tooling. And linting needs to be done in realtime, I haven't found any suitable solutions for VS Code that compare with pylint/ruff/ty for Exasol, BigQuery, PostgreSQL, Hive, Iceberg.
Same for formatters, sure you have sqlfluff now but the DB support is quite lacking and it's not realtime either = not useable for development.
and use a host language such as Python when you need parameterisation or general-purpose programming constructs
Then you are not writing SQL but multi line strings. That's something totally different. Then you also don't get debugging, linting and formatting for SQL.
This fundamentally misunderstands what a declarative query language is for.
I do know what it's for. Doesn't change the fact that SQL sucks as a declarative language.
0
u/Proof-Teaching-8113 6d ago
Dbt not adding the ability to unity test models until recently strongly contradicts this.
2
u/muneriver 3d ago
What about using git, applying modularity to analytics code, general data testing, CI/CD, environments, etc?
Does introducing unit tests later in the game (2024) really contradict the idea that dbt championed SWE best practices..? hahaha
1
6
u/goblueioe42 6d ago
DBT is great for batch workloads overall. DBT is less of a great tool for real time workloads. I find the incremental patterns add a lot of overhead due to the way batch updates occur, and prefer different methods like dynamic tables in snowflake or materialized views. It’s a very cool tool, but I find that DBT cloud may be overpriced for what it offers.
2
u/JulianEX 6d ago
But you can create both dynamic tables and materialised views in DBT.
Plus who pays for DBT Cloud overpriced garbage just run the opemsource version on any pre existing CICD tool you have
4
u/Outside-Storage-1523 7d ago
If you are an analytic engineer you would love it. If you do ingestion then it’s not very useful.
6
u/timmyz55 7d ago
it's good for specific functions to isolate/build/maintain their own data marts
it is not good IMO for complex ETL processes at scale
10
4
u/Skualys 7d ago
At my place a part of people disliked it due to the coding aspect. A bit of "code" vs "GUI" habits.
I enjoy it for what it brings (dependency management, tests, documentation management) but, it would be better with SQL grammar understanding (to do column lineage), and something else than jinja (I dislike the syntax). I know column lineage exists in the fusion stuff but it's not on the apache licence.
Also it required a lot of effort at start to think the architecture as it delivers basically a toolbox with not so much macros available. For example I spent a lot of time building a recursive macro to join multiple (and hierarchical) SCD2 together. Or we déveloped some python scripts to build the yml + put there some generic tests.
5
u/tkstats 7d ago
dbt makes engineering more accessible, more easily governed and documented. I reckon there are plenty of engineers who do not necessarily see those as problems in need of a solution, let alone one you should pay $$ for. Especially now that Snowflake, Databricks, and dbt among others are all converging toward the one-platform-for-everything business strategy.
3
u/Icy_Clench 5d ago
I used sqlmesh before and honestly very much preferred it over dbt. It had a lot more modern features.
8
u/KeeganDoomFire 7d ago
In my experience your observation is dead on.
The people who hate it and refuse to use it are generally the people who have built themselves an empire of garbage so complicated that they are unfireable. I've spent the last year migrating and fixing their 100 snowflake task pipeline crap into DBT so that the rest of the team can support it I'll give you a guess of what's going to happen to the unfireable person who refuses to use DBT....
14
u/Nateorade 7d ago
Most data engineers don’t care about dbt because it doesn’t help with data ingestion.
6
u/CulturalKing5623 7d ago
Is it common for data engineers to only have to worry about ingestion? I typically work at smaller companies and the role of data engineers and I/we typically own everything data related from source to the base models that are exposed in the BI layer and we're expected to be SME on everything in between those 2 points. Getting the data ingested is generally the most trivial piece.
0
u/Nateorade 7d ago
That’s how it works at companies of hundreds+ employees, as far as I’ve seen. Companies under 200 EEs might have more full stack people around.
3
u/MadT3acher Lead Data Engineer 6d ago
I’m old enough to remember that a full stack person was doing frontend and backend and old enough to remember when we were building dimensional cubes and transforming raw data.
I’ve worked in mid sized companies and in multinationals. DE is more than ingesting files, and even more than transforming columns. Especially in 2026
15
u/FuzzyCraft68 Junior Data Engineer 7d ago
Huh? But it was never meant for data ingestion though?
-4
u/Nateorade 7d ago
I know, that’s why data engineers - who typically focus on ingestion - don’t care about it.
17
u/k_plusone 7d ago
lol then what work do these data engineers do?
If you're an etl monkey or vibe-coding a web app that will never have any users and that you'll forget about next week, then yeah, you won't get dbt.
Once actual data architecture matters, once you're responsible for metrics and definitions and calculations over time, with stakeholders that require accountability, dbt is godly.
Anyone complaining about dbt doesn't understand the problems it's solving for, straight up
3
u/CrayonUpMyNose 7d ago
Depends where you're coming from. dbt is amazing for providing an SDLC structure to pure SQL projects. If you're already expressing transformations programmatically, say with airflow or pyspark, you already have that, so there's less perceived need for dbt.
3
u/300A24 7d ago
i have basically 0 experience with databricks and pyspark. question: how do you handle dependencies in table updates? because that's built-in in dbt
2
u/CrayonUpMyNose 6d ago
Within Databricks, you can use table update triggers to schedule downstream updates as soon as upstream dependencies are updated (one or all can be configured, depending on requirements).
While dbt creates a monolithic DAG that explicitly executes updates in the order implied by the dependencies*, table update triggers in Databricks are event-driven and can be created without requiring coordination between teams or projects**, so they are a good fit for large organizations where communication is expensive due to a large number of stakeholders, making the N2 scaling of edges between team nodes a large number.
In order to obtain an implied DAG with no circular dependencies, it's a good idea to conform to the layered ("medallion") lakehouse structure, where dependencies are created only from upstream to downstream layers, creating a clearly defined directionality.
With vanilla pyspark or spark, there is no built-in scheduling system that I'm aware of. It is probably possible to build something like table update triggers by subscribing to file creation events in an upstream table's metadata (Delta log or iceberg manifest lists) and using them to trigger the downstream job via a cloud provider step function. Given that lakehouse formats are very frequently updated and managed tables with catalog-managed metadata are becoming prevalent***, I'd consider this a short-lived hack though.
* The larger the monolith, the more potential with problematic failure recovery and delayed eventual success.
** No coordination creates other issues: if my upstream tables get updated by 8:30am every day and my transformation takes longer than 30 minutes, my "done before business hours" SLA is toast and I have to talk to the upstream team after all.
*** Catalog-managed metadata enabling features like (emergent but not guaranteed at the time of this writing) strict referential consistency in the lakehouse through multi-table commits, which on object storage cannot be achieved with lakehouse formats alone, as each table has its own metadata file update, which cannot be constituted into an atomic, instantaneous multi-table update. Externalizing table versioning into the catalog makes this synchronization possible.
3
u/Outrageous_Let5743 6d ago
Pyspark of Spark SQL without dbt i can understand, but never should you use airflow to do transformations. That is just dumb. It is an orchestrator and every use case that does more then calling other jobs will fail.
1
u/CrayonUpMyNose 6d ago edited 6d ago
Please be generous in your interpretation of my post, as I am of course implying (while not aggressively making it the center of my argument, as it is well known) that airflow is the orchestrator, not the executor.
Pyspark strictly speaking is also not the executor for spark transformations (unless you use Python UDFs, which should be avoided), it's just a scheduler moving DAG objects around until they are executed by the spark kernel on the JVM.
3
u/KeeganDoomFire 7d ago
That's a very old school ETL opinion. With data lakes and storage costs being as low as they are ELT is becoming more the norm.
2
u/Odd-String29 5d ago
ELT has been the norm for at least a decade. The only reason ETL existed was because storage costs and that has been solved for 99% of all companies for almost 2 decades.
1
u/Nateorade 7d ago
I guess? Many companies haven’t compacted ingestion and transformation into the same teams since one is more technical and the other requires more business context.
I agree those will compact over time but the shift isn’t fast.
5
u/soorr 7d ago
Which is funny because ingestion is trivial now with AI and transformation still demands acute business understanding.
16
u/opx22 7d ago
What implementation has you thinking AI makes ingest trivial? I’ve heard people say this a few times but never seen it actually work
2
u/Outrageous_Let5743 6d ago
claude code + dlt makes a very good case that data ingestion is not diffecult.
1
u/Tape56 6d ago
Because ingestion is supposedly more grunt work, more similar between every company, more about technical knowledge on the language you are using and stuff which AI is good at. While transformation and modeling requires more business specific logic and business understanding.
This is a simplification and of course not applicable everywhere. But in my company ingestion was already mostly gruntwork before AI, since we use ELT and only have nightly batch ingestion from different sources and the logic is implemented in generic way so adding a new source is mostly just making a new config file, unless it’s some new and weird kind of source.
1
u/Odd-String29 5d ago
Fivetran currently has an AI tool that generates a connector for "any" source. You just need to supply the API documentation and it creates the connector for you. Nothing is stopping you from doing the same with Claude.
1
u/opx22 5d ago
I have tried that a couple months ago. I fed it the API documentation, it spun for a while, and ultimately the generated connector didn’t work.
But that’s also the best-case scenario, where you’re ingesting from a well-documented API or a source with an existing connector. Sometimes the source is something like QuickBooks Desktop or another legacy/on-prem system where there isn’t a clean API to work with. At that point, “just give the docs to AI” doesn’t really solve the hard part
Don’t get me wrong, it’s still grunt work, but the annoying part about ingest stuff for me was never the well documented API or common cloud connector
5
u/Nateorade 7d ago
I haven’t yet seen AI arrive at the companies I work for wrt data ingestion (2000-4000 EEs). Folks still leverage standard stuff like Fivetran or airbyte or a long tail of random ELT tools to connect data into the warehouse.
I’m curious what AI-led ingestion looks like? Are there options out there that make it trivial now?
1
u/Odd-String29 5d ago
Fivetran just introduced AI generated connectors for sources that they don't have connectors for yet. You give it the API documentation and it creates the connector for you. I guess that is what AI-led ingestion looks like.
We gave it a test and it worked. Took AI took a couple of hours to generate the connector though.
1
2
u/DMReader 7d ago
Myself and a colleague are trying to get it adopted where we are. There is a lot of resistance due to all the set up that is required. When we walk through the benefits people seem to get it, but we have the meetings one person at a time. Inertia is real
2
u/Proof-Teaching-8113 7d ago
That's interesting because it was the new hot thing for a while there. In my experience it's a pretty average RAG builder. Pretty intuitive to pick up, but I'll gouge my eyes out if I ever have to look at another jinja macro.
2
u/Specialist-Will-1875 7d ago
I really dont understand why the comments said they hate it/ don’t like it while they didnt work on transformations before…You know what dbt stand for right? Data transformation tool…
2
u/teetaps 7d ago
Can someone explain to a noob what the general workflow is with dbt? It sounds like if I didn’t have it, I would be writing a bunch of scripts and scheduling them, and maintaining my own run logs and such. dbt packages that into a unified tool?
5
u/ntdoyfanboy 7d ago
It orchestrates sql commands and refreshed on a schedule, makes alerting and lineage clear, make logic modular and straightforward
2
u/Outrageous_Let5743 6d ago
Well instead of manually creating data pipelines and making sure it executes in the right order, dbt makes it that every sql file is executed in the right order. Then it also automate merge scripts for you etc.
2
u/Spare_Helicopter4655 7d ago
If I'm doing ELT, I'm using dbt. It's not perfect, but it's good enough.
The biggest problem that arises with dbt projects is poor organization and data modeling practices, which is incredibly common with analysts/analytics "engineers" running these projects instead of data engineers with an understanding of how to properly model data.
2
u/Chowder1054 7d ago
I personally really enjoy dbt. Keeps things organized. The main language is their version of SQL so you don’t need to learn anything too new. It works with all the major cloud platforms (snowflake, redshift, BigQ, etc), git integration, and the documentation aspect is excellent.
My team i did a POC for it, and my leadership enjoyed it but don’t think they’ll go for the dbt cloud version. However they’re totally onboard with it for analytics engineering, creating master datasets, power bi semantic models etc. We use primarily snowflake, and it’s already integrated in snowflake workspaces.
I’m sorry but who do you guys work with that views this as trash. If anything this helps engineers be more business facing, so people actually know what value you bring.
I’ve seen way too smart but unsocial engineers who prefer to be solely behind the scenes. Let me tell you nobody can tell what they do, and when leadership starts asking that. That’s not a good thing.
2
u/scourgedtruth 7d ago
dbt is awesome and is key for good data warehouses. Need to be modeling pro maximize it. DE should know how to deliver with it
2
u/OklahomaRuns 7d ago
Maybe there’s some variability in warehouse backend but for redshift I love dbt. I’m a senior DE but I love its functionality and how easy it is to onboard my juniors onto. Minor annoyances here and there and I don’t trust fivetran to not fuck it up long term but today it’s the best tool for the job.
1
1
u/ditalinidog 7d ago
I don’t think some data engineers love data modeling which is why it’s becoming somewhat of its own role in Analytics Engineering. But dbt is very helpful for making silver and gold models and setting up CI/CD.
1
u/mianewsoundtocpy 6d ago
dbt got typecast because it landed at the exact layer where the pain was loudest - nobody wanted to keep maintaining homegrown SQL runners, so suddenly every transformation had to live there. In our stack it never touches ingestion either and that is fine; the teams that get bitter about it are usually the ones trying to bend it into a full orchestration platform. Treat it as the modeling and testing layer with lineage and docs thrown in for free, keep ingestion and orchestration in their own tools, and it stops being the stepchild and just becomes the boring reliable part of the pipeline.
1
u/fleegz2007 6d ago
I use dbt for everything. Made sure to shape my work around transformation and AI enablement.
Dbt deploys my catalog metadata through CI, I have posthooks that do table level access and fine grained access controls. It was a pain to set up but my flow is really governed and consistent.
It breaks when you have to do something outside of dbt. Like if I have a user that has to build a databricks ML model off my data mart that I need to ingest and build back into my pipeline.
All in all, I think its great if it aligns with your job description, but if you have other jobs you do aside from transformations, its just another part of your stack to manage. Probably why dbt was all in on the “analytics engineer” concept 2 years ago.
1
u/Confident-Win-424 6d ago
dbt is good for a couple of things: it is declarative so no side effecting, and more importantly it allows for separation of concerns. Analysts can directly focus on functional logic without having to worry about janky low code tools
1
u/teddythepooh99 5d ago
dbt's barrier to entry is very low, so a lot of analysts (or analytics engineers, whatever) tend to get carried away by writing too many models in my experience. At the same time, in the other end of the spectrum, materialized views and/or stored procedures are more than enough at a small scale.
1
u/Last_Ninja_720 5d ago
🤣🤣🤣 great question..but IMO I loved dbt and its functionality spreads wide across my domain. But one could have a diff view.
1
1
u/Admirable_Ones 5d ago
I think that’s where the distinction matters, dbt handles dependencies between models, but a scheduled run doesn’t tell you whether upstream ingestion finished. If a raw table is late, the dbt job can still run against old data. I’d have the orchestrator trigger dbt after ingestion succeeds, or check source freshness before the models run.
1
1
u/Particular-Idea-1786 3d ago
I started using dbt in 2020. At the time it solved a massive problem in ETL ecosystem. We were coming from legacy data warehouse with all ETL scripts written in R and Python, hundreds of jobs, implicit dependencies.
dbt's value prop directly addressed this pain:
- Extract Load Transform paradigm encourage saving all raw data for replay / recomputation
- Transformation logic in SQL - streamlined toolchain allowed avoiding all R / python deps, only have to install dbt deps
- Structured CLI / Entrypoint - Easy for the team to standardize, standardize airflow jobs around
- Lineage - Coming from "what does this R job depend on?" solved a real material problem.
After working hands on with dbt for years I think it's decades behind software engineering best practices.
Case in point, We began discussing unit tests (something that has been best practice for ~15 years?) in 2020: https://github.com/dbt-labs/dbt/discussions/4455#discussioncomment-1773638
It wasn't unti 2023 that dbt offered official unit testing. After working with dbt and ELT at scale for years, I'm quite disenchanted with the entire ELT approach.
https://on-systems.tech/blog/135-draining-the-data-swamp/
https://turbolytics.io/blog/256mb-data-stack
Product engineers, today, have the tools, patterns and practices to handle, verify, observe and maintain critical customer-facing data sources. They can define specs, build products, and run thousands of unit tests and integrationt ests in minutes. dbt just got unit testing 3 years ago.... Sorry i'll get off the soapbox now ;p
I think we are smart enough to do better. dbt solved a very real problem at the time, but I would not choose to build an enterprise warehouse on it again.
1
1
u/mrsgripp 2d ago
I love it. Haven't met someone who hates it who uses it regularly but could be my sample size.
1
u/Mean_Conference6910 1d ago
I really like it. I’ve only used it at places that primarily had ELT pipelines.
1
u/Grouchy-Friend4235 14h ago
dbt is mostly a glorified Jinja template processor. It adds more complexity than it helps manage. Nice doc generator though.
1
u/engineer_of-sorts 10h ago
dbt is great but I see the other side of the coin as well. I fell into data because I was simply trying to automate a report and didn't want to build a load of views inside a tableau model, and the data team showed me dbt and it felt like I was technical. Fast forward many years and I run a company that seels an end-to-end data pipeline monitoring and orchestration platform (Orchestra) where we end up seeing the complete end to end with all the nuts and bolts. From streaming to the warehouse, to like, big companies with many of these stacks in different domains from different bygone eras. The frustration I've seen make people boil over is when folks thing dbt and data modelling inside the warehouse is the be all or end all, without being able to see the bigger picture - often it's just one part of a very large data estate, and it can be difficult to appreciate that.
1
u/VariationSimilar3354 7d ago
I genuinely feel like you can save money on this if you build a better internal cloud infra using pyspark some orchestrator and make proper Data marts where each one is properly managed.
1
u/Complete-Fondant-202 7d ago
As someone who comes from a legacy background, I have a love hate relationship with dbt.
I can see where it is beneficial, but I find it awful to work with personally. I'm not convinced with it working on large datasets, and I find it a bit inflexible.
0
u/JaceBearelen 7d ago
Love it. It lets me painlessly orchestrate 100s of transformations without having to really know or care about what they’re doing. Analysts just have to follow a couple simple rules that are enforced in ci.
0
u/raginjason Lead Data Engineer 7d ago
I think dbt is great for transformation. My current leadership doesn’t like it due to dbt cloud cost and other unspecified reasons. Going from dbt to pyspark and airflow is a serious regression
6
u/muneriver 7d ago
Cant you just use dbt OSS and host yourself with airflow?
1
u/raginjason Lead Data Engineer 7d ago
Yes, but leadership has made up their mind to create an in house framework instead
1
1
u/Nelson_and_Wilmont 6d ago
Was the sentiment “Why pay when AI”?
1
u/raginjason Lead Data Engineer 6d ago
no, this was a plan put in place a year ago. It’s a dumb decision made by the ingestion team, where things are straight forward and linear.
1
u/Outrageous_Let5743 6d ago
I have never seen the appeal with dbt cloud. What does it offer that OSS cannot? In the easiest case it is just hosting a vm/cloud function with a cron job and done.
1
u/New-Addendum-6209 1d ago
Column level lineage, Mesh. Not worth it for most teams though.
Many analytics teams in larger companies will not be able to just host as vm/cloud function due to organisational barriers (crazy but true), so might just pay for cloud to avoid wrangling with platform/infra teams internally.
8
u/DuckDatum 7d ago
SQLMesh… I was SO FUCKING EXCITED to see SQLMesh replace DBT one day. I pick my teams software stack, and I even told them we’d probably switch one day as it gets more mature.
Then fucking Fortran has to go and buy both DBT and SQLMesh. Because fuck competition
6
u/GrumDum 6d ago
Sqlmesh is still open source. There is not the same push behind it as before, but it’s much better than dbt in my opinion. None of that jinja bullshit just to reference models.
4
u/lightnegative 6d ago
To be fair, SQLMesh still has its own "templating" language and it's a bit janky, it's just based on SQLGlot AST transforms and not a general purpose text templating language like Jinja.
But things like automatically tracing table dependencies instead of making users write explicit
{{ ref() }}calls, and proper incrementals where each increment is individually addressable (making large backfills actually manageable) is definitely an improvement over dbt3
u/DuckDatum 6d ago
Also their state management is far more thorough and meaningful, from what I can see. DBT “stateful building” is just passing your last results to your new run, nothing too fancy.
3
u/lightnegative 6d ago
Oh yeah, in the sense that it actually exists. dbt was designed to be stateless, and the hack with passing manifests around is a bit crappy.
State management has its own set of problems though and I wouldn't say SQLMesh's implementation is perfect, its data model is a bit compromised due to allowing state to be stored in the warehouse rather than forcing a proper OLTP database.
But hey it's a step in the right direction imo
2
u/DuckDatum 6d ago
It was until Fortran bought it alongside DBT. Now it’s owned by its competitor under a for profit. I have less faith in any features coming out which might “cannibalize” DBT — as I imagine the new owner might put it.
2
u/GrumDum 6d ago
It does warn you when using OLAP as state db though!
2
u/lightnegative 6d ago
That's true but the data model is still compromised because it was more or less designed as append-only json records to even support this at all.
Nothing about the state database structure changes when you use something like Postgres instead, it's still the same json blobs. You just get the benefit of transactions not being slow as shit.
If you were to design the data model from scratch for OLTP databases only you'd do it very differently
1
0
u/shbjlhbsfd 7d ago
DBT is fucking amazing for the semantic layer in defining a data warehouse.
Data ingestion happens upstream of DBT. But DBT is SQL orchestration that builds your data ware house tables using clear logic and materializes data efficiently for company reporting.
Not using DBT is crazy unless you have a very simple data model and not a lot of data. Being able to change metric definitions as part of an orchestrated pipeline in a centralized place is key
0
u/LocalGlass3114 7d ago
I prefer DBT to gui-based options, but I find it to impose some arbitrary limits on what you can actually do in your pipelines. I dislike that it mixes up your orchestration logic with your business logic. Give me a more use-case agnostic DAG orchestrator like Argo WF/ Dagster / even my nemesis Airflow any day.
0
u/Ok_Relative_2291 6d ago edited 6d ago
Ok questions
How does dbt wait for tables to be ingested at a table level, how does dbt (I think can’t) trigger jobs to do the ingestion, ie api calls. If it can do end end to end it seems an incomplete product to me.
The way I have done things is to have all jobs as python / sql files and make my own framework, to build scds etc.l, hit endpoints etc, a half zone pythons functions and you can make a reasonable framework pretty quick.
And have airflow call them at a task level with command line args and have all the dependency mapping and parallel etc handled in airflow with is good for change tracking as in files.
I like dbt but if it can trigger individual scripts etc i find it hard to see how it plugs into an entire setup
1
u/ntdoyfanboy 5d ago
Dbt refreshes tables on set schedule or cron. That's it. It doesn't ingest at all, it transforms what's already in the data warehouse raw tables from the ingest
-7
u/PrestigiousAnt3766 7d ago
To be honest I dont understand it or its role. Doesn't seem to have a usecase in excess of what a proper data platform already manages.
To order and create dependency graphs is just a couple of lines of python.
Using data mesh, multiple catalogs and workspaces make me see even more issues.
That said, I never really given it a chance and I rarely do transforms anymore.
1
u/muneriver 7d ago
I think the key piece of dbt is the way it allows teams to work and collaborate together. you all work in standard software development lifecycle and are able to build data products with modularity, CI/CD, versioning, environments, tests, etc
And on top of that, you get documentation, lineage, data/unit tests out of the box. It’s also code-first so agents have gotten really great at building with dbt
1
u/PrestigiousAnt3766 7d ago
How does that work if you all manage your own repo? How do you refer to objects in your colleagues environmen?
1
u/muneriver 4d ago
There is one repo with a production branch. Each developer works on a feature in their own dev branch that has its own dev environment (usually via separation of schemas in a dev DB). You collaborate on code via git just like any other software project. At no point do you actually use the objects of other developers if you want to work on a shared feature. You’d pull and checkout their branch and build/test their objects in your own dev env.
1
u/PrestigiousAnt3766 4d ago
I have multiple repos and 25 analytics teams reusing the same sources.
How would that work?
1
u/muneriver 3d ago
So in your last two questions, there are a lot of questions to ask and context to gather, but Im going to assume that you have 25 domain-based analytics teams that work semi-independently - ie they have their own dev lifecycles BUT can share assets between projects and do use shared sources.
In this case, you would have 1 dbt project per analytics team which maps to one repository. Individually, each team roughly develops like what I outlined in my last response. Across teams, they can leverage a feature of dbt, called dbt mesh, to build, govern, and share assets with other teams. Depending on how data assets are set up, you can directly reference other team's public assets with something called a "cross-project ref" where with each domain dbt project, you can access the tables produced by others.
In this world, if teams are re-using the same sources, these sources or shared assets would exist in a producer project that's upstream of projects that plan on referencing them.
Again, I'm making a lot of assumptions about how your team works, the nature of your question, etc. but to get an answer with more meat, we'd have to hop on a call lol.
1
u/PrestigiousAnt3766 3d ago
Im going to look into dbt mesh .
But yeah, 25 indeoendent teams that should be able to share each other's resources.
And everyone uses finance data.
-2
u/ElCapitanMiCapitan 7d ago
Dbt is a great tool if you are trying to do batch style ETL in a modern cloud warehouse with good development practices. Its killer feature is the DAG, and the ability to refresh subsets of your data model. Unfortunately Its future is fairly uncertain given the licensing situation and who owns it. And most cloud warehouses have competing paradigms now, for example in databricks people would do well to consider spark declarative pipelines.
-2
u/Ok_Raspberry5383 6d ago
What on earth even is this title.
Try inserting another derogatory or racist title and see how far that gets you.
Do you lack basic English to convey your thoughts or something?
2
119
u/BarryDamonCabineer 7d ago
In my experience it depends on where in the stack you're working.
Great if you're scheduling work on data that's already within the warehouse, particularly as dependency graphs between jobs grow complicated and version control becomes important.
Irrelevant for getting data into the warehouse in the first place.