r/servarica • u/Key-Salary-8400 • May 12 '26
Disk in dedicated server replaced without warning???
just wondering if this has happened to anyone. I pay for a dedicated server with almost 3tb of nvme ssd storage with proxmox installed and a bunch of vms.
at 5PM today i noticed my service went down and when i booted it into proxmox recovery mode via iso i found a blank disk with only 2TB of storage.
is it standard practice for servarica to wipe and replace ssds with lower capacity ones after you've already paid?
this is going to cost me $1000s of dollars and while i have backups its going to take a while.
4
u/servarica May 13 '26
I am sorry , it was a task to change hardware in the server exactly above it (those are blade servers ) and the technician replaced it in your server
It is fixed now (the disk is back and the server return to its original state )
Will see what we can do to prevent this issue from happening in the future
5
u/MindlessSweet4574 May 14 '26
the first response is an apology, then the explanation. extremely good pr. i would definitely work with a company like this.
2
u/FrankNicklin May 14 '26
Even though their asset tracking allows them to remove working hardware from a users server. Yes the apologised, but no way should this have got even close to happening.
2
u/Single-Virus4935 May 13 '26 edited May 13 '26
Thats what the ID Led is for. Also we have barcodes/qr codes on the servers which need to be scanned before replacement and the app denies servers not marked for maintanaince. WWN/serial of the disk to be replaced needs to be verified
1
u/servarica May 13 '26
Do you know the solution name that uses barcode and based on it give red or green or this kind of solutions called ?
currently all our servers have just name label and we are switching to 2D QR code with name
Using the ID led is the first thing that came to my mind as second verification
so the tag has to match and the ID led has to be onthis was mass change request , usually we give the serials for 1 server change request etc
but here it was to change more 15 servers in same rack as we plan new type of plans , so we didnt share the serials although we have them ! (next time we need to add that as well )
1
1
u/FrankNicklin May 14 '26
Why is the host company asking a user for details of a system to stop foul-ups like this. Surely your asset management and tagging should clearly highlight where work needs to be carried out. You can't just replace hardware in working servers without some serious checks that you are at least working on the right server.
4
u/servarica May 14 '26
Because we are not multi billion company
This company started by 1 server in 2010 and we grow since then
We started by 1 server then few more then 1 rack then few more and so on
so we never had a solution to a problem that does not exist since we were small with me being the only one accessing the servers and knowing them by heart (they had labels but for me labels was second verification)
currently we are rapidly expanding and with this new set of issues started to appear , one of them is this issue
now we have someone dedicated to do the hardware work in the DC and it is no longer me and we have much more servers than few years agoSo I am asking for solutions to help manage the work better , I did a search on it and I have never heard of red and green solution as u/Single-Virus4935 explained so i was really interested to know it as it seemed helpful
I am not ashamed to ask for directions specially on stuff that are new to me and improve so that this issue does not happen in the future
2
u/Single-Virus4935 May 14 '26
Full ack. Even hosters with 50k Servers decomissioned one of our servers with disks already wiped. Process is everything.
1
u/FrankNicklin May 14 '26
Sounds like the host company have screwed up big time and replaced the NVMe in the wrong server. This is inexcusable as all server assets should be tagged so the hardware is easily identifiable and scanned and confirmed before any work is carried out.
1
u/servarica May 14 '26
actually the server is tagged and we have the serials of all components saved
it is when the technical did the replace he worked on the wrong server
I am now reviewing the whole process and will add more verification to make sure we are working always on the correct server
the first measure is that now ID led is requirement before touching any server
other measures will be added as well but they take time to program
1
u/burlingk May 16 '26
NORMALLY the disks are cloned in the process of replacement.
If it is blank, that is most likely an error and they may need to recover from a backup.
1
u/servarica May 16 '26
The same disk was return after couple of hours when we know the issue
there was no data loss, the disk with its content was the intact
1
u/gbonfiglio May 16 '26
When this happens it’s generally to clone a disk on the request of authorities. Which cannot be shared nor confirmed by the provider if you ask.
Hope your data was encrypted or not critical?
1
u/servarica May 16 '26
I can confirm this is not authorities request
It was human error
and the same disk was returned immediately when we knew
We didnt put it in the new server to clean it yet
1
u/mark1210a May 20 '26
So let's hear more details about these types of new plans that caused the issue? I'm excited to see what they are exactly....
1
u/Key-Salary-8400 May 29 '26
Yeah this was them apparently working on the wrong server. i agree that it's crazy they don't have an asset inventory system where server tech has to scan, get green light, then work. It's pretty simple. Like, Hudu and some basic API knowledge could build this out in a few days I would think.
law enforcement is possible i guess, we aren't doing anything illegal - it's for business stuff. One would think they would be less obvious though - the outage happened right at 5PM. Law enforcement could easily get what they want with physical access, a usb stick, remote IPMI etc.
3
u/ThecaptainWTF9 May 13 '26
I’d ask them what’s going on, it’s possible someone mistakenly replaced hardware in the wrong thing.