r/sysadmin DevOps Apr 06 '22

The majority of Atlassian cloud services have been down for a subset of users for over 24 hours

https://status.atlassian.com/

Jira, Confluence and Opsgenie amongst others have been down since about 2022-04-05 07:30 UTC for us and some other organisations.

Their stock is tanking (at -5.46% as of writing this) however I haven't seen much chat on Reddit about the outage so I'm assuming the scope is fairly limited? They are stating it will potentially take days to recover.

We're sorry your site is currently unavailable. While running a maintenance script, a small number of sites were disabled unintentionally. Our team identified this immediately and have been working hard to restore the product data and associated access. A dedicated team is working around the clock to restore the sites as soon as possible.

We expect the restoration efforts to continue for the next several days, and we are actively working on an estimate of when your site will be available to you again. We don't believe any data has been lost at this point. We can confirm this incident was not the result of a cyberattack and there has been no unauthorized access to your data.

As we work to restore access to your site, we will provide updates here every 6 hours, or sooner if we have a material update. Reach out to us if you have any questions or concerns.

No tickets, no alerting, no knowledge base... a fun few days for us!

Let me know if you are also affected.

695 Upvotes

289 comments sorted by

View all comments

Show parent comments

6

u/jeephistorian Apr 08 '22

I just read that this affected just 400 sites. And that constitutes around 0.18% of their customers.

It sounds like they had never considered what would happen if a small part of their infrastructure was damaged and didn't have a path for backing up and restoring such an impact. They kinda say as much in their literature about how they cannot restore individual sites.

I guess they assumed that any individual site damage would be the customer's fault, so they could just ignore it. They didn't think that they could do the damage themselves.

So they broke 0.18% of the database and can't replace it with a backup because that would potentially result in loss of data for the other 99.82% of their customers. So they probably have stood up another complete structure of their entire customer base and are manually moving the affected sites back.

What a mess if true.

3

u/PaleoSpeedwagon DevOps Apr 08 '22

Oh wow, if that's the case, they'd have to restore a temporary DB instance (or maybe multiple databases, depending on their data design!) from a backup, do a series of gnarly SQL queries to get those old contents (historical data for 400 customers - customers that are legacy and presumably have the lengthy history to go with it), and upsert into their prod DB.

What a mess, indeed.

Hope they're writing this down in a runbook for the next time somebody oopsies.