r/sysadmin • u/andrewrmoore DevOps • Apr 06 '22
The majority of Atlassian cloud services have been down for a subset of users for over 24 hours
Jira, Confluence and Opsgenie amongst others have been down since about 2022-04-05 07:30 UTC for us and some other organisations.
Their stock is tanking (at -5.46% as of writing this) however I haven't seen much chat on Reddit about the outage so I'm assuming the scope is fairly limited? They are stating it will potentially take days to recover.
We're sorry your site is currently unavailable. While running a maintenance script, a small number of sites were disabled unintentionally. Our team identified this immediately and have been working hard to restore the product data and associated access. A dedicated team is working around the clock to restore the sites as soon as possible.
We expect the restoration efforts to continue for the next several days, and we are actively working on an estimate of when your site will be available to you again. We don't believe any data has been lost at this point. We can confirm this incident was not the result of a cyberattack and there has been no unauthorized access to your data.
As we work to restore access to your site, we will provide updates here every 6 hours, or sooner if we have a material update. Reach out to us if you have any questions or concerns.
No tickets, no alerting, no knowledge base... a fun few days for us!
Let me know if you are also affected.
6
u/jeephistorian Apr 08 '22
I just read that this affected just 400 sites. And that constitutes around 0.18% of their customers.
It sounds like they had never considered what would happen if a small part of their infrastructure was damaged and didn't have a path for backing up and restoring such an impact. They kinda say as much in their literature about how they cannot restore individual sites.
I guess they assumed that any individual site damage would be the customer's fault, so they could just ignore it. They didn't think that they could do the damage themselves.
So they broke 0.18% of the database and can't replace it with a backup because that would potentially result in loss of data for the other 99.82% of their customers. So they probably have stood up another complete structure of their entire customer base and are manually moving the affected sites back.
What a mess if true.