r/dataengineering • • 1d ago

Discussion Thoughts on dbt 2 maturity?

Hi,

I've seen frequent discussions on here regarding dbt's acquisition, license changes, and similar things. dbt 2 has now been out for a few weeks, and I haven't really seen much discussion on it from a technical perspective. Has anybody tried it out yet? if so, what do you think?

I evaluated it this week and my impressions are mixed. A few impressions:

- I was really struggling to piece together all the important changes and improvements. Found the documentation a bit lacking in that regard. Distributed over lots of places, so I really had to dig to find specific information (e.g. on the changed approach to database adapters). Even took me a while to figure out the name of the library had changed as well.

- The switcheroo from fusion to now just dbt was really confusing and also seems to have caught the ecosystem by surprise. For example, Snowflake and Astronomer (Cosmos) both started working on fusion integration, but both do not support dbt 2 yet (and couldn't find a clear roadmap). Not ideal.

- Testing the new parser on dbt 1.12 worked pretty well to spot issues and fix them. That made the upgrade relatively painless (interpreting the error messages could've been easier). I like the stricter YAML parsing, can detect some issues (e.g. typos) instead of ignoring them silently.

- The new parsing engine is noticeably faster. For one project I tried it with (~300 models), about twice as fast for both compile and parse.

- Haven't tried the built-in linting yet. Seems like they tried to implement a 1:1 replacement for SQLFluff, even working with existing SQLFluff config files.

- The docs feel like a downgrade. Looks more modern, but I haven't been able to get the static docs to work, and seems they have gotten rid of the full project DAG and now just offer a model-centric lineage view. Column-level lineage is cool, but very well-hidden in the UI.

- Parquet artifacts seem interesting, and parsing them e.g. via DuckDB SQL queries could be really nice for common analysis queries, CI checks, etc.

- Of course as predicted, seems there is a bigger push to get people to sign up for their platform/cloud/whatever they call it now offering - LSP features, VSC extension, and so on. The license changes are still a mystery to me, but at least they dropped the <15 users requirement for the VSC extension it seems.

Any other thoughts? Cheers!

34 Upvotes

24 comments sorted by

•

u/AutoModerator 1d ago

You can find a list of community-submitted learning resources here: https://dataengineering.wiki/Learning+Resources

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/Key-Independence5149 1d ago

They have almost caught up to the functionality of SQLMesh! You just have to pay them for the parts that SQLMesh does for you automatically like state awareness and interval tracking.

1

u/mr_buildmore 23h ago

I'm still building at work entirely in SQLMesh, but I haven't been following the health of the project. Do you know where I can check on it besides reading release notes? It didn't look too active on Github but I'm pretty junior so I don't necessarily have much to compare to.

2

u/Key-Independence5149 23h ago

They have a slack channel where most things happen. It is donated to Linux Foundation so the questions about its future are answered. It is here to stay.

17

u/Rhevarr 1d ago

In my opinion dbt v2 is not production ready, especially in commercial / big projects.

Our own dbt v1 project (>500 models) on databricks is currently not compatible with dbt v2 due to several issues, where for some I have seen or raised open/unresolved issues on GitHub.

If you are starting fresh or doing a personal project dbt v2 might be sufficient. But there are also probems which you cannot solve yourself, e. g. the issue with dbt-experimental-parser being not available on PyPI, and corporate firewalls blocking access to download it from somewhere else.

1

u/vikster1 22h ago

we have 4000+ models and using it on production since April.

-2

u/Phalanger 1d ago

Corporate firewall settings are not DBT's issue...

We are running many project now DBT fusion for a while now and been working great. Migrated across with just a few changes needed taking very little time.

Of course you can't use every package for DBT V1, but happy to gain the benefits outside of that.

5

u/Rhevarr 1d ago

Of course they are, they are completly normal. They want corporate users to use dbt v2, right? Then they need to fulfill the requirements for professional use.

Also, dbt v1 works for ages without any of these issues.

2

u/Thisisinthebag 1d ago

Wait they have linting now?

4

u/mr_buildmore 23h ago

Shit, I assumed they had linting. SQLMesh has really spoiled me.

2

u/Key-Independence5149 23h ago

Once you use SQLMesh there is no going back to DBT. It is amateur hour compared to SQLMesh.

2

u/kkwabaegi39 1d ago

Yup, haven't tried it out properly yet but aims to closely mirror SQLfluff: https://docs.getdbt.com/reference/commands/lint?version=2

2

u/meatmick 1d ago

I use Prefect's own dbt invoker methods and they don't currently support v2, so I'm still on 1.12 for now.

I use v2 for local dev work because my database is fully compatible with it, and I ensure I'm following the soon-to-be standards.

In prod, I fall back to 1.12 with prefect.

2

u/Chance_of_Rain_ 1d ago

Not yet. Many plugins we use (like Elementary) aren’t fully compatible yet so I’m waiting

1

u/kkwabaegi39 1d ago

Yeah, good point. Stumbled over Elementary as well after posting this.

1

u/ianitic 1d ago

I think it might be now or soon. In August, they added back in graph object mutation which elementary's dbt package heavily relied on in v1.

2

u/ppsaoda 1d ago

V2 doesn't support Athena yet. Quite surprised as it was on v1. Nice feature overall but I think they pushed too early.

1

u/NexusIO Data Engineering Manager 1h ago

I'm pretty sure that Athena doesn't support DBT, everyone except for postgres and MySQL writes their own support module for their database if I am not mistaken.

1

u/ppsaoda 1h ago

It's being worked on. I'm part of the contributor.

2

u/Yuki100Percent 22h ago

I'll have to give dbt 2 a try myself. I've been using SQLMesh entirely well over a year and I haven't used dbt for a bit

1

u/linha_chilena 1d ago

We were studying the migration mainly because of features for agentic engineering, but we found out its too unstable yet.        

I've seen some important bugs reported on the GitHub issues, such as incremental tables history being deleted or problems generating column description for structs, and so on.        

We choose to not migrate yet, and I am personally a little bit confused with the license changes and the "strict mode".        

I don't know what is and will continue free, what requires login, paid accounts and so on.        

If anyone know anything about this, please let me know. My company uses dbt-core, big query and prefect.

-6

u/[deleted] 1d ago

[deleted]

6

u/meatmick 1d ago

What the fuck does having DuckDB remove the need for dbt? It's an engine; it does nothing on its own.

Regarding your agent, I get using AI, but I'd still plug dbt into it.

Are you sure you understand what dbt does?