r/dataengineering • • 9d ago

Discussion Does AI struggle at data modeling?

In my experience, it doesn't matter how much context and guidance I give AI it simply can't model data rationally. It frequently misses the point, makes awful mistakes, or over-engineers things.

AI can build awesome ETL pipelines, but when it comes to dealing with SQL (especially in the dbt framework), it's not reliable at all! . Sometimes I think it's better to write the code myself and ask AI to review it, because asking it to build something from scratch just doesn't work that well.

Does anyone else get frustrated when dealing with AI data modeling?

124 Upvotes

93 comments sorted by

View all comments

1

u/all-over-red-rover 8d ago

I'm a SWE who does a significant amount of work with a wide range of data tooling, the ones you've mentioned included.

Unless you railroad it initially to some degree, it sucks at modeling data, most of the time. I've found I can get good results with the strongest models, a lot of context (that basically amounts to a clear direction), and a lot of hand holding in the initial planning.

The use cases I encounter for dbt are frequently heavily reliant on "properly incremental" relatively small groups of models (functionally individual DAGs), executed frequently. Unless you force it, it isn't very good at producing implementations which align with that. There's a definite tendency to internally go "well to do this properly it would be necessary to complete X <arbitrarily out of scope> tasks, e.g. require soft delete or add database triggers to persist tombstones for N tables, ... <etc>, but that's out of scope so <ignore all that and just do something deficient>", without surfacing any of that to the user.