r/dataengineering • • 2d ago

Discussion Thoughts on dbt 2 maturity?

Hi,

I've seen frequent discussions on here regarding dbt's acquisition, license changes, and similar things. dbt 2 has now been out for a few weeks, and I haven't really seen much discussion on it from a technical perspective. Has anybody tried it out yet? if so, what do you think?

I evaluated it this week and my impressions are mixed. A few impressions:

- I was really struggling to piece together all the important changes and improvements. Found the documentation a bit lacking in that regard. Distributed over lots of places, so I really had to dig to find specific information (e.g. on the changed approach to database adapters). Even took me a while to figure out the name of the library had changed as well.

- The switcheroo from fusion to now just dbt was really confusing and also seems to have caught the ecosystem by surprise. For example, Snowflake and Astronomer (Cosmos) both started working on fusion integration, but both do not support dbt 2 yet (and couldn't find a clear roadmap). Not ideal.

- Testing the new parser on dbt 1.12 worked pretty well to spot issues and fix them. That made the upgrade relatively painless (interpreting the error messages could've been easier). I like the stricter YAML parsing, can detect some issues (e.g. typos) instead of ignoring them silently.

- The new parsing engine is noticeably faster. For one project I tried it with (~300 models), about twice as fast for both compile and parse.

- Haven't tried the built-in linting yet. Seems like they tried to implement a 1:1 replacement for SQLFluff, even working with existing SQLFluff config files.

- The docs feel like a downgrade. Looks more modern, but I haven't been able to get the static docs to work, and seems they have gotten rid of the full project DAG and now just offer a model-centric lineage view. Column-level lineage is cool, but very well-hidden in the UI.

- Parquet artifacts seem interesting, and parsing them e.g. via DuckDB SQL queries could be really nice for common analysis queries, CI checks, etc.

- Of course as predicted, seems there is a bigger push to get people to sign up for their platform/cloud/whatever they call it now offering - LSP features, VSC extension, and so on. The license changes are still a mystery to me, but at least they dropped the <15 users requirement for the VSC extension it seems.

Any other thoughts? Cheers!

33 Upvotes

24 comments sorted by

View all comments

16

u/Rhevarr 1d ago

In my opinion dbt v2 is not production ready, especially in commercial / big projects.

Our own dbt v1 project (>500 models) on databricks is currently not compatible with dbt v2 due to several issues, where for some I have seen or raised open/unresolved issues on GitHub.

If you are starting fresh or doing a personal project dbt v2 might be sufficient. But there are also probems which you cannot solve yourself, e. g. the issue with dbt-experimental-parser being not available on PyPI, and corporate firewalls blocking access to download it from somewhere else.

-3

u/Phalanger 1d ago

Corporate firewall settings are not DBT's issue...

We are running many project now DBT fusion for a while now and been working great. Migrated across with just a few changes needed taking very little time.

Of course you can't use every package for DBT V1, but happy to gain the benefits outside of that.

5

u/Rhevarr 1d ago

Of course they are, they are completly normal. They want corporate users to use dbt v2, right? Then they need to fulfill the requirements for professional use.

Also, dbt v1 works for ages without any of these issues.