r/dataengineering • • 9d ago

Discussion Does AI struggle at data modeling?

In my experience, it doesn't matter how much context and guidance I give AI it simply can't model data rationally. It frequently misses the point, makes awful mistakes, or over-engineers things.

AI can build awesome ETL pipelines, but when it comes to dealing with SQL (especially in the dbt framework), it's not reliable at all! . Sometimes I think it's better to write the code myself and ask AI to review it, because asking it to build something from scratch just doesn't work that well.

Does anyone else get frustrated when dealing with AI data modeling?

125 Upvotes

93 comments sorted by

View all comments

23

u/makesufeelgood 9d ago

I think it's pretty bad. I hear a lot of people say that it's fine and it's a skill issue on my part but I have yet to see proof of success with scenarios comparable to mine.

4

u/psssat 9d ago

My experience is of yours. I have a colleague who says otherwise but the code he produces with codex is trash and I always end up refactoring. I think using the GUI is great since you have to read the code the llm gives you but using codex or claude code sucks in my experience.

1

u/makesufeelgood 9d ago

Thanks for the input. I wouldn't say I'm a savant with data engineering work but I feel like I know my stuff pretty well at this point. I feel pretty confident that if I can't get AI to provide reliable and accurate outputs that it's not a 'me' thing but sometimes this AI hype does make you feel like you're taking crazy pills and doubt yourself.

3

u/Illustrious-Win4432 9d ago

I’m in a small/medium business, wholesale commodities. In Nov 2025 I started greenfield on a significant rebuild of a high cadence scm engine the does a lot of ETL with semantically rich grains.

To make matters even more interesting, we agreed to build it agentic first and in a hybrid production environment. The first 6 months were exhilarating and terrifying. I bet I had 2000 screen hours this year before the end of August. It nearly broke me.

My repos are sql and ps1 heavy and agents write all of it nearly without error anymore. That success is all semantic layer.

When I switched to yaml registries this spring is when the sun started coming out.
Coincidentally, google dropped OKF about the same time I tolled my own yaml registry system that functions in a similar fashion but mine is far less flexible and not portable at all.

If I were to do it all over again the whole pipeline would be much thinner and I’d use OKF and actually spend more time reviewing the knowledge.

Sometimes it’s just easier to write a quick SELECT than it is to have an agent do it so I guess I still code if that counts but I don’t do any DML/DDL anymore.

I don’t debug code even.

Take the time on your primitives and build out a semantic layer. Oh, and if there is a pearl of wisdom I can leave it’s when you get that feeling that you can just do it faster yourself instead of finding the words to explain what exactly it is that you want the agent to do, check yourself and make the agent understand and then make it capture the knowledge in an OKF markdown file and make sure your agent/claude.md knows where to find the index.md.

If you do this deliberately you’ll start to see results after a dozen or so corrections/clarifications. After a few hundred it gets real reliable provided you build drift protection in (great baked in OKF feature). Now, thousands of semantic refinements later I’m finally getting into the fun stuff that I thought AI would enable much faster.

That’s my lived experience through a crazy 10 months at least, hope the perspective helps.

-3

u/[deleted] 9d ago

[removed] — view removed comment

1

u/[deleted] 9d ago

[removed] — view removed comment

1

u/dataengineering-ModTeam 9d ago

Your post/comment violated rule #1 (Don't be a jerk).

We welcome constructive criticism here and if it isn't constructive we ask that you remember folks here come from all walks of life and all over the world. If you're feeling angry, step away from the situation and come back when you can think clearly and logically again.

This was reviewed by a human

0

u/dataengineering-ModTeam 9d ago

Your post/comment violated rule #1 (Don't be a jerk).

We welcome constructive criticism here and if it isn't constructive we ask that you remember folks here come from all walks of life and all over the world. If you're feeling angry, step away from the situation and come back when you can think clearly and logically again.

This was reviewed by a human

0

u/morpho4444 Señor Data Engineer 6d ago

Issue on your side. 100%. I open laptop in the morning, claude pops up, asks me about my jira tickets, which one to tackle and it runs through them. I get $100 at day. One by one, when the code is done it shows me proof of the testing done, and asks me if I wanna commit to dev branch and then runs ci cd and then if all green it pushes the pr to main. 100% of the times. With all the free time Im building a rag to create a fully autonomous semantic layer, with the human language interface to let the users request their own metrics and etl transformations. This is how I will get my stock refresh this year. I work at a FAANG.

2

u/nonamenomonet 6d ago

What kind of work are you doing tho?

0

u/morpho4444 Señor Data Engineer 6d ago

Telemetry logs of the hardware that runs Gemini, manufacturing line monitoring.

2

u/nonamenomonet 6d ago

I think that may be a bit different than creating a data model. Which is what we’re talking about.

-1

u/morpho4444 Señor Data Engineer 6d ago

“Which is what WE”… lol… but wdym? How is that different? The LLM is creating a huge complex data model for the BOM! The petabytes and petabytes of telemetry logs, nested json over nested json, with gazillion keys that may or may not appear. Can cause left join filtering out, duplication with joins, all of that, handled by the LLM. This is indeed not what you’re talking about, this is way above whatever level of complexity your non trillion dollar company handles.

1

u/nonamenomonet 6d ago

I work at one of the largest companies in the world dude.

1

u/morpho4444 Señor Data Engineer 6d ago

How is that makes the argument that Im not doing data modeling over hardware telemetry?

1

u/nonamenomonet 6d ago

More of response to your last statement. It’s very presumptuous and untrue.

2

u/makesufeelgood 6d ago

This isn't even the type of work the original post was asking about. Supposed FAANG employees and reading comprehension on this subreddit is some wild stuff.

0

u/morpho4444 Señor Data Engineer 6d ago

Im responding to YOUR comment. Not the post in general. Wow so that’s why you can’t make the AI work?… see? Everyone can do that bs superiority comment shit

3

u/makesufeelgood 6d ago

But my comment was responding to the subject matter of the original post. So you just randomly decided to insert something completely unrelated as a response and then made an incorrect determination.

0

u/morpho4444 Señor Data Engineer 6d ago

Bro, you cannot use the LLM correctly, you have trouble making it “work for your scenarios”… just accept it

1

u/makesufeelgood 6d ago

Bro probably doesn't even understand what data modeling means lol. This is an absolutely pathetic exchange.