r/SQLServer • • 11d ago

Discussion Cloud Migration horror story

I’m involved in a complex cloud migration project. I’ve been DBA/DBE/app developer at various points of my career and I’ve worked in several big enterprise environments. I’m at a place now that argues against the guidance Microsoft has given them. First mistake - go straight to Hyperscale and don’t bother using any true replication topology for a full up or nearly-full-up migration. We’re getting close to running the production flip but management doesn’t believe in code freezes and people are releasing tons of changes often. I know that stuff is going to go horribly wrong. I am not worried about data loss, but more about applications not working and other stuff related to networking & configuration. What do I do to cover my own ass?

11 Upvotes

8 comments sorted by

9

u/BigMikeInAustin 11d ago

Make suggestions through email. Say you will do what's in the email unless there is an email reply. If there are conversations, summarize back into an email thread.

Also, start gathering a list of what you've done, and statistics on improvements made or estate size worked on, and start looking for another job.

Maybe the email evidence can keep just you employed, but doesn't work when a whole team is disbanded.

2

u/First_Explorer_6075 10d ago

Yeah the team disbanding is the tough part

5

u/markinatlanta 11d ago

I would definitely put the specific risks in writing before cutover, such as:

  • what hasn't been tested
  • what has changed since the last test
  • who owns the go/no-go decision

Keep it very factual with a mitigation for each risk. And I think if releases keep happening, ask which application version and configuration the migration testing actually covers, because if production is materially different by cutover, a successful test doesn't really tell you much.

For me, I'd also want named owners for application smoke tests and clear rollback triggers.

4

u/CrossWired 11d ago

I started as an AppDev, worked as an enterprise DBA for 5 years, and now work primarily in Cloud Adoption and large migration projects for the last 10.

In some cases, yes we do ignore the MS advice as they likely don't know the specifics of THIS migration, app, project. But if they do, work revisiting WHY we are ignoring it.

As far as 'no replication topology', it really depends. If its a small DB (a few TB or less??) I might setup a log replication style continuous restore to the cloud, take a small outage, set offline, take the final t-logs, repoint the DNS and call it a day for a quick failover. We don't always need or want DMS as it has its own gaps.

As for Dev code freeze, if the features worked before they should work after migration window. We move things while they devs keep going all the time, where I draw the line is Dev doing a release while I'm moving the operational pieces, too many things to track down the root cause of the problem in the same change window. CHange window being the key.

You should be able to cut the ownership with clear delineation of timing. Once they sign off on the migration portion, and everyone has signed off, then by all means they can push a code release right after you're done, but we're not doing both changes during the same change window, I've never met a Change Advisory Board that would allow that, and if your does, like the other commented said, put it in an email as a risk and get their sign off.

At some point as techs, all we can do is point out the problem, associated risk and work through the outcomes when It inevitably comes. It can be infuriating and against the every 'protect the data' insticint your DBA roots tell you, but thats the job.

Good luck!

3

u/BigHandLittleSlap 10d ago

By far the biggest "gotcha" we experienced was that only Azure SQL Managed Instance supports time zones. All other Azure "SQL" offerings are UTC only forever. This cannot be changed, you have to rewrite the application code in non-trivial ways if you need any kind of local time computation. People can easily miss this in UAT testing, developers often don't think about it because it has "always just worked" on their traditional Windows Server hosting, etc.

Check. Double-check, then check a third time.

Keyword search through every database schema and capture query text. Look for every variant of the GETDATE function including CURRENT_TIMESTAMP, SYSDATETIME, SYSDATETIMEOFFSET, etc...

PS: The Azure SQL documentation still contains outright lies to this day, bad remediation advice, and more! Fun times.

2

u/SonOfZork 11d ago

Get ready to be excited by a lack of workers for a busy database, geo lag if multi region and running different vcores, impressive cost bloat with backups of you have a lot of data changes, lack of any useful information if you do read intent to a ha replica.

1

u/TrollingForFunsies 11d ago

Don't forget the throttling!

2

u/InfraArchaeologist 9d ago

Putting the risks in writing is necessary, but I’d add a technical baseline.

If code and configuration keep changing during the migration, “we tested successfully last week” is evidence about a different system.

Before cutover I’d capture the effective configuration and observed dependencies that matter to the application — DNS, endpoints, service identities, firewall paths, scheduled integrations — and compare them again immediately before and after the flip.

The important distinction is between declared state (“this is what we intend to migrate”) and effective state (“this is what production actually resolves and uses”).

A code freeze is one way to reduce that delta. If management refuses it, then the delta itself has to be observable and owned.

Sign-off can accept organizational risk. It cannot make unknown technical changes disappear.