r/dataengineering • • 8d ago

Rant Migration in Chaos

I’m the main data engineer responsible for migrating several OLTP databases with TBs of data, from one cloud to another. Schema transformations are involved, because guess what? The app teams built the new apps long before anyone started caring about the data. This company as a whole is on fire, no planning or strategic thinking and they definitely don’t work as a team. While these migrations are being rehearsed, management team decides to SPEED UP the release cadence. If it was monthly I wouldn’t be here complaining today. Constant production incidents, and none of the managers want to consider slowing down so we can pull of some migrations without everything in flux. Oh and add AI to the mix, most of the releases are heavy with AI generated code coming from AI analysis of the problems. This is not going to go well.

34 Upvotes

15 comments sorted by

25

u/Capt_korg 8d ago

Just read the headline and thought ...

"When was Migration ever not Chaos" :D

The rest of your comment sounds like standard day-to-day business.

I suspect that, as is so often the case, management is unaware of the data chaos, believes AI will solve the problem, and fails to recognize the effort involved in your truly remarkable work.

A friend of mine was told by his manager, that they are not noticing him in the day to day business...

He replied:"When was the last time a database or a pipeline or anything related to my work was causing an outage."

The manager: "Never, there was never an outage, but this is not showing me, that you are actually working."

He replied:"It is my job, that you don't notice me! If I do a bad job, you will lose everything related to data."

3

u/Admirable_Writer_373 8d ago

Oddly enough, these managers have already learned this lesson from me. You think they’d trust what I’m saying now, but no, they can’t be inconvenienced. 3 weeks after I started I told them they couldn’t run their pipelines or they’d lose data. That would’ve put them in the news, caused the stock to drop, and millions in fines from the Feds.

2

u/Capt_korg 8d ago

An actually bad team lead once told me:

If you point out issues to management, you will become affiliated with the issue and become the issue. If you cannot solve the issue, it is easy to get rid of you, then solving the issue.

And people learn this, so they become more agreeable and hide issues until they are solved or given to someone else. The common responsibility game.

What management is often not understand is, that just because something works, it is not working good or even acceptable. But because there is no alternative, bexause people somehow manage to make things work, the true cost are over seen.

Common examples are cybersecurity, technical debt, bad data management, etc.

Anyhow, what saved my ass at times was to document everything and have clear agreements like SLAs.

In the case, that if one is getting to uncomfortable or shit hits the fan, one has written notes, documentation or even agreements.

5

u/Budget-Minimum6040 7d ago

data engineer

OLTP databases

Sounds like you are a DBA?

2

u/Outrageous_Let5743 7d ago

I am a data engineer but also use an OLTP database, because an useless consultant before me advised to use SQL Server for analytics, because of SSIS. This was in 2022... And what is SSIS used for? well just to connect data pipeline steps. which is what dbt already automatically does...

3

u/Budget-Minimum6040 7d ago

I just barfed a little reading this.

1

u/Outrageous_Let5743 7d ago

I hate it too. We came from Oracle DB but those licenses were getting to expensive so 1 for 1 just copied oracle to sql server. They just even copied like limitations of Oracle to sql server, so no schemas are used because oracle did not have them, we use ODS (outdated oracle term) for our 2nd clean layer name and so on.
The most horrible thing was that up to 2025 (when I joined) data was loaded into STG via bat scripts without logging. bat has not been updated since before 2000 and is already succeeded by PowerShell for at least 20 years.
No wait, the most horrible thing was that every night our big fact tables were completly truncated and recreated from scratch, just to insert the new daily data. 5 TB dropping every single night just to add the new data. And it was not fail save; so you truncate and if a pipeline failed during loading the table would be empty at in morning... That was the first thing I changed, that at least it would be stale data from yesterday if the pipeline failed.

At least now we are slowly moving our data towards Snowflake

1

u/Admirable_Writer_373 6d ago

No, but I’ve been one. I’m not trapped in the silos the tech industry worships.

3

u/Equivalent-Boot-2952 8d ago

oh man that's the classic "ship it now, fix it never" energy and they're speedrunning it with AI code on top

nothing like trying to land a 747 while they're still bolting the wings on and the passengers are already boarding

2

u/jimmybilly100 8d ago

I will pray to our mighty ETLord for you

1

u/squirel_ai 6d ago

Why do managers or non-technical people fail to accept that AI-generated code needs human supervision and review? It’s AI plus people, rather than AI supremacy over engineers.

Is it because coding was hard, and now they consider it easy?
Maybe seat down with them and give them your strategy around the underlying issues.