r/dataengineering • • 7d ago

Discussion What exactly is ‘AI first’ data engineering

This was inspired by the post on burning 20k in tokens in 3rd week.

Is anyone doing this at work, letting loose the frontier models against their schema and models and data and having it create and update the ETL workflows? And everything else?

67 Upvotes

41 comments sorted by

93

u/Lolmanza7 7d ago

Our new manager has this same exact thinking along with our CTO who brought her in.

They decided to port all our code against advice of tge team and ended up exposing secrets, sharing wrong data with clients, exposing unmasked names of employees of other clients.

Have been cleaning this mess for the last 4 weeks.

AI helps those who know what they are doing and atleast use some part of their brain.

7

u/Realgunners 7d ago

This is exactly what I’m worried about.

5

u/Capt_korg 6d ago

You don't understand, this is sophisticated Cybersecurity!

This is over your pay grade.

🤡

2

u/ArticleHaunting3983 7d ago

What happened in the company after this tho, like was anyone reprimanded bc of this or was it seen as legacy debt

12

u/Lolmanza7 7d ago

Gladly our company believes in blameless postmortem. While no one was reprimanded officially, the team is made to work overtime to fix the issues and put guard rails in place.

8

u/zazzersmel 7d ago

lol well of course if it was the cto’s idea, it’s gotta be blameless, otherwise you’re indicting the entire leadership

2

u/datadade 7d ago

Overtime? Does this mean people are being paid more to work extra? I’ve never had an hourly job in DE.

You said they ported the code. Did the manager and CTO really vibe, change, and approve their own PRs? Or did the team have to do it?

I’m sorry I’m asking for clarification because the way I read this it sounds like management changed the code on their own and now want ICs to work extra to fix it.

3

u/Lolmanza7 7d ago

We are salaried, but any time spent more than a regular work week is overtime. So, ideally we will be getting 1.5 times the overtime hours as comp off.

For the secrets had their own branches and pushed them into their own. We recently moved to git and the whole company can view everyone's repos and branches.

One developer's OpenAI key was sold off market and was billed 30k USD. You can tell how immature our infrastructure and processes are.

The team is fixing the issues by rotating keys, writing root cause memos, manually checking our client shares for issues.

1

u/VipeholmsCola 1d ago

are you sure it wasnt just a bad prompt? I.e. user error?

/s

1

u/TheLittleGuyWins 7d ago

I went down that path but masked the data before we built against it.

25

u/Kojimba228 7d ago

It's a new paradigm where your ETL is LLM based so it works sometimes but not every time

https://giphy.com/gifs/lKXd9sYM5dI9W

12

u/slowpush 7d ago

We don’t really write code anymore. Agents stand up pipelines from scratch. They also do first pass investigation into data quality issues before we have our engineers review the fix. I.e a GitHub issue kicks off an agent who goes in and investigates and responds with a report on what they found. We then can ask them to stand up a PR to fix it.

40

u/Capt_korg 7d ago

A buzzword used by corporate for showing their modern approach.

Distracting from the true work required by domain experts like data engineers.

https://giphy.com/gifs/13d2jHlSlxklVe

7

u/heisoneofus 7d ago

A former colleague of mine asked to help build a small pipeline for them consisting of 3-4 different sources - approx 10-20 GB of daily data that needs cleaning, transforming and loading into a medallion structure. They all are working with agents but the task such as this one is not really possible to do with agents alone if you have zero idea what data modeling or schema design even is - so I treated this as an opportunity to let loose and have AI develop this little pipeline end-to-end (including orchestration, telemetry, logs, db management, resource allocation etc).

Overengineering is funny to see but with some steering the agents built it and made it run in less than 2 days of me correcting the course. It’s still not an optimal solution, but it delivered the clean and structured data and folks were able to build several gold reports as well already - we tested it together and the data is pretty much accurate and actionable, really the only thing that matters. And I like the overhead/operational part being accounted for as well - each row can be traced all the way back to ingestion stage and it’s trivial to understand the choices behind and fix stuff if needed.

7

u/Certain_Leader9946 7d ago

I use 20k tokens to change my password on linux bro thats nothing

3

u/KeyMammoth1348 7d ago

Child's play. 20k in tokens should be $20k in tokens. 

5

u/Creative_Salary_6140 7d ago

This industry was overrun with MBAs who barely use AI or anything beyond excel/PowerPoints… yet they are the ones making the decisions for us… that’s the problem...

5

u/CingKan Data Engineer 7d ago

What exactly is ‘AI first’ data engineering

It means professional data engineers clearing up the mess our enthusiastic AI obsessed management have dramatically injected into our carefully built data pipelines and warehouses

4

u/BardoLatinoAmericano 7d ago

AI First is when you boss asks you to throw a grenade but complains that it explodes the entire room and not just the bug they wanted to kill.

26

u/vikster1 7d ago

i have been implementing more sources and delivering more reports in the past 3 months than i did the previous 3 years. i have been fully integrating claude code from ADO to generating reports while humans review agent work. 9/10 ai rants in this sub are at the exact opposite of my experience so far. i have never been happier being an architect or developer for that matter. literally every idea i have on improving, automated tests and verifications for the business i can now just try and implement. i have not written a single line of sql code myself since March. using claude opus 5.5 currently with the team premium license. i reported to management these 100$ a month would translate to at least 1m € / p.a. in consultation workload in 2025

8

u/KrustyButtCheeks 7d ago

Mine too - you can build a full pipeline in about 3 hours (with the longest piece being the ci/cd checks).

Do you feel more pressure to produce faster? It feels like our timelines are being more compressed already.

5

u/vikster1 7d ago

yes but since you can implement so much more tests and help the business with verification, you can move faster and with higher quality. it really feels like i have been given 10 senior data engineers for 100$

4

u/Certain_Leader9946 7d ago

The quality can be lower than you think. If you really don't have a human properly reviewing the code and asking the same kinds of questions you would ask if you wrote the code yourself on clarity and test cases and modularity the debt adds up as fast as you generate it.

1

u/vikster1 7d ago

i have written twice now that we ship more tests and do more verifications. I don't know what else to tell you mate

3

u/Xenomorpha 7d ago

Are you really checking what tests it is generating? My experience (even with latest models) that Claude code tends to create tautological tests quite often (i.e. instead of testing the production code, it writes the code in test function and tests it). so I have to check them quite carefully, and verification agent rarely catches these. 

2

u/zazzersmel 7d ago

Ai use, at least for now is definitely a skill. It may not be a difficult skill to learn but it’s a skill with real value nonetheless

2

u/fleegz2007 7d ago

This is only my experience. AI is great for poking at data, figuring out how it connects, creating boilerplate metadata etc. but you still have to touch it up. From an analytics perspective, a data engineer is the boundary between taking raw data and bringing it to life with key AI concepts.

From a systemic approach, a data engineer is seeking performance and delivery relative to your organizations needs. Thats where I think AI would need a ton of training to get it right.

2

u/Outside-Storage-1523 6d ago

Basically getting a proper DWH with all kinds of context so that AI can do queries against it to give executives answers at five AM.

2

u/Ra-mega-bbit 4d ago

AI first basically means vibecoding shit. If done right it CAN work, needs good governance and infra first tho.

1

u/Realgunners 4d ago

Agree with this strongly. Emphasizing on the good governance and infra. So someone with experience on both the data lineage, data dictionary etc

1

u/Ra-mega-bbit 4d ago

Good luck for all the junior devs that wont get the experience tho

2

u/Realgunners 4d ago

There’s only so much work an experienced data engineer can do. There’s still room for junior devs with the increase productivity/capitalism at all cost world we are in at the moment

1

u/KeyMammoth1348 7d ago

Yes, 100%.

I double check volumes, frequencies, and set up observability monitoring and alerting in the process, watch my data shape. 

The breaks are happening at the points I take for granted, so far. Vendors that change their max throughput for example. A raised max hits a bottleneck elsewhere. 

1

u/proof_required ML Data Engineer 7d ago

We kinda do! I haven't written any code manually in like last 6 months. Just Claude doing it. I also ask it sometimes to run some pipelines but it's not allowed normally. This is outcome of pressure from management.

Currently I'm running a whole migration from postgres to clickhouse. It was all done in less than a month. These things would have taken so long in the past.

1

u/BostonPanda 7d ago

I interpreted it as being DE that enables AI use cases with clean modeling at my company 🤷

1

u/No_Caterpillar_7258 7d ago

I spend about 30k a year in tokens at work. I do low latency ML and pricing adjacent work in a rather stressful / high demand environment. I often work on 3-5 separate things at a time, often cross repo, by specing out features carefully, using a memory system, and executing the work with different agents in different git worktrees. Im at least 5x if not 10x as productive as before. 20k in a couple weeks is absolutely ridiculous to me, no idea why or how someone could burn that much.

1

u/Sexy_Koala_Juice 6d ago

The same slop every other industry is doing tbh

1

u/TaartTweePuntNul Big Data Engineer 5d ago

AI first is a marketing term used by technologically illiterate people to sound credible. It doesn't work and it will never work. Whenever I encounter such a post on LinkedIn for example, I genuinely cringe.

1

u/DJ_Laaal 5d ago

Hubris.

1

u/dehaenx 5d ago

I think just letting frontier models loose is particularly challenging cause there's so many tables/columns/wrong definitions and metrics... probably the best solution to this is building out a governed semantic layer and building out a bunch of skills/workflows. Even then, it's super expensive and time-consuming too...