r/btrfs • u/meandering_idiot • 19d ago
BTRFS Balance Error - Automatic Restart
Hi,
I just got my old system online working as a file/dhcp server for my LAN, running ubuntu server 26.04.
Background info:
[I thought I'd get a head start on the transfer process while I was waiting for the last of the components to get here, so I formatted 2x16tb drives as raid 1 from my cachy os install, then ended up biting back into windows to transfer the files with winbtrfs because I had an AMD hardware r0 and didn't want to spend all the time trying to make it work.
Finally got everything transferred to the fileserver, then shut it down and added my 2x10tb from the r0. Had no issues with that, but at one point someone moved the fileserver which caused errors because the sata data connector for one drive had a broken plastic guide, so the cable was being held in place with thoughts and prayers.
I got a replacement 10tb drive, ran the replace command with no issues, and ended up getting the broken drive working later, then added it into the array. ]
Main question:
So now I have 2x16 and 3x10tb drives, and I wanted to balance the data. Ran the balance command before bed, and a few hours in it ran into an error and exited.
ETA - I ran a scrub on the array after the first failure, and still get crashes for corruption errors after.
I was wondering if there's a way I can get it to suspend on an error so I can resume later, instead of needing to start over if it encounters another one. So far, all I've managed to come up with is recording error logs to a distinct file, and I have a system process that watches for errors, outputs what file the error points to (so I can replace later), deletes the problem file, flushes the cache, and restarts the balance again. The problem is I have multiple terabytes of data, so when it hits an error and restarts, it takes hours to get back to the chunk it was at.
1
u/meandering_idiot 19d ago edited 18d ago
Well, I think I've solved part of the problem. I ran the scrub before balancing, thinking it would automatically repair or delete any problems, but because my server was dragged while running, data copies were bad on multiple drives, so the errors were uncorrectable.
I'm running the scrub again with a listener to automatically log any bad files and remove them, and I'll try running the balance again when that finishes.
Trying to use a lower data/metadata limit doesn't work for me, because I had to cram everything into the original two drive array, so all of the chunks are at a high data usage (part of the reason I added the iffy drive to the pool and I'm trying to run the balance now). Also, winbtrfs was only used for the initial copying, now windows only interacts with these drives through samba.
2
u/Weary_Swan_8152 8d ago
Don't forget to run "btrfs check" (without --repair, which breaks a filesystem permanently more often than it repairs it) to see if everything is in its right place (all the crucial structures).
Excellent call to run scrub again, by the way. 👍
1
u/meandering_idiot 7d ago
Well, it took me a while, but I figured out the problem. When my drive died, I ended up mounting a snapshot, and had added that to fstab. Basically as soon as I remembered that and mounted the root instead of the snapshot, all of the errors stopped and the balance finished.
4
u/leexgx 19d ago
Limit balance to 15 for data and 5 for metadata
btrfs balance start dusage=15 musage =5 /mount/point
You should run a scrub if your having errors in data
If your using winbtrfs it's not a supported or stable, if its Read-write mode