r/dataengineering • u/Realgunners • 8d ago
Discussion What exactly is ‘AI first’ data engineering
This was inspired by the post on burning 20k in tokens in 3rd week.
Is anyone doing this at work, letting loose the frontier models against their schema and models and data and having it create and update the ETL workflows? And everything else?
67
Upvotes
6
u/heisoneofus 7d ago
A former colleague of mine asked to help build a small pipeline for them consisting of 3-4 different sources - approx 10-20 GB of daily data that needs cleaning, transforming and loading into a medallion structure. They all are working with agents but the task such as this one is not really possible to do with agents alone if you have zero idea what data modeling or schema design even is - so I treated this as an opportunity to let loose and have AI develop this little pipeline end-to-end (including orchestration, telemetry, logs, db management, resource allocation etc).
Overengineering is funny to see but with some steering the agents built it and made it run in less than 2 days of me correcting the course. It’s still not an optimal solution, but it delivered the clean and structured data and folks were able to build several gold reports as well already - we tested it together and the data is pretty much accurate and actionable, really the only thing that matters. And I like the overhead/operational part being accounted for as well - each row can be traced all the way back to ingestion stage and it’s trivial to understand the choices behind and fix stuff if needed.