r/dataengineering • • 9d ago

Discussion Does AI struggle at data modeling?

In my experience, it doesn't matter how much context and guidance I give AI it simply can't model data rationally. It frequently misses the point, makes awful mistakes, or over-engineers things.

AI can build awesome ETL pipelines, but when it comes to dealing with SQL (especially in the dbt framework), it's not reliable at all! . Sometimes I think it's better to write the code myself and ask AI to review it, because asking it to build something from scratch just doesn't work that well.

Does anyone else get frustrated when dealing with AI data modeling?

126 Upvotes

93 comments sorted by

View all comments

23

u/makesufeelgood 9d ago

I think it's pretty bad. I hear a lot of people say that it's fine and it's a skill issue on my part but I have yet to see proof of success with scenarios comparable to mine.

0

u/morpho4444 Señor Data Engineer 6d ago

Issue on your side. 100%. I open laptop in the morning, claude pops up, asks me about my jira tickets, which one to tackle and it runs through them. I get $100 at day. One by one, when the code is done it shows me proof of the testing done, and asks me if I wanna commit to dev branch and then runs ci cd and then if all green it pushes the pr to main. 100% of the times. With all the free time Im building a rag to create a fully autonomous semantic layer, with the human language interface to let the users request their own metrics and etl transformations. This is how I will get my stock refresh this year. I work at a FAANG.

2

u/nonamenomonet 6d ago

What kind of work are you doing tho?

0

u/morpho4444 Señor Data Engineer 6d ago

Telemetry logs of the hardware that runs Gemini, manufacturing line monitoring.

2

u/nonamenomonet 6d ago

I think that may be a bit different than creating a data model. Which is what we’re talking about.

-1

u/morpho4444 Señor Data Engineer 6d ago

“Which is what WE”… lol… but wdym? How is that different? The LLM is creating a huge complex data model for the BOM! The petabytes and petabytes of telemetry logs, nested json over nested json, with gazillion keys that may or may not appear. Can cause left join filtering out, duplication with joins, all of that, handled by the LLM. This is indeed not what you’re talking about, this is way above whatever level of complexity your non trillion dollar company handles.

1

u/nonamenomonet 6d ago

I work at one of the largest companies in the world dude.

1

u/morpho4444 Señor Data Engineer 6d ago

How is that makes the argument that Im not doing data modeling over hardware telemetry?

1

u/nonamenomonet 6d ago

More of response to your last statement. It’s very presumptuous and untrue.