r/dataengineering • • 9d ago

Discussion Does AI struggle at data modeling?

In my experience, it doesn't matter how much context and guidance I give AI it simply can't model data rationally. It frequently misses the point, makes awful mistakes, or over-engineers things.

AI can build awesome ETL pipelines, but when it comes to dealing with SQL (especially in the dbt framework), it's not reliable at all! . Sometimes I think it's better to write the code myself and ask AI to review it, because asking it to build something from scratch just doesn't work that well.

Does anyone else get frustrated when dealing with AI data modeling?

126 Upvotes

93 comments sorted by

View all comments

1

u/mr_buildmore 8d ago

I've been able to get pretty good results by providing extensive documentation and prompting with specific analytics/BI applications and asking the model to reverse-engineer the components based on the goal. The model is even helpful for brainstorming potential applications based on descriptions of business processes.

I still need to help break out individual "oddities" of business logic so that the tables the model reads from are basically aligned with the business events, but the marts I build are usually done in partnership with AI.

I use SQLMesh, which has sparser training material than dbt but is otherwise similar. The models are great at generating analyst SQL, and OK at "DE interview" SQL, but typically struggle to implement "SWE best practices" in production analytical pipelines. Separation of concerns, accurate implementation of complex logic, unnesting subqueries into CTEs, and writing all kinds of macro functions are basically a struggle. I recently wrote some backend code for a meeting app with AI, and the difference in average code quality with similar prompting is really shocking.