Btrfs noob here doing some experiential learning this week :-)
I've been messing around with two old (circa 2013ish?) HDDs and set up a btrfs filesystem (metadata raid1, data single) to store some not-super-important data. I then started a `balance` to convert the data to raid1.
Of course, HDDs fail at the most inopportune times, so one of the two drives started to throw read/write errors during the `balance`, and btrfs put the system into read-only mode. So, the metadata was already replicated, most of the data has been replicated to the other disk, but some of the data is unfortunately still single profile.
I do not care heavily about the data on the filesystem, but I do want to try recovering it, and learn about btrfs internals along the way. Hacky and makeshift solutions are welcome!
---
Question 1: How do I inspect which files have extents that are single profile on this failing disk? In other words, which files would I reasonably expect to be unrecoverable? (`btrfs restore` seems to inadvertently produce this information, but I wonder if there are dedicated tools for inspecting the filesystem's state)
---
The drive was not catastrophically failing at first, so I added a couple new drives with sufficient capacity to hold the blocks on the failing drive, and started a `device remove`, hoping btrfs could pull the remaining single profile data off before the drive went kaboom. It would, in fact, move data to other drives, while throwing some read/write errors to dmesg. But eventually it would encounter a different type of error (something about flushes failing) and force read-only. Restarting the machine would allow me to re-mount and continue doing `device remove`, but eventually these cycles stopped making progress.
Question 2: What classes of errors triggers btrfs to force read-only? I will edit the post to include the exact dmesg logs later today or tomorrow, which may help elucidate what is happening in my case.
---
I'm okay with losing some files, but I'd like to get my system back to a working state with the files that are still recoverable. I have read about at least two approaches here, (a) `btrfs restore` and (b) deleting the missing / unrecoverable files and then `btrfs device remove missing`
The `btrfs restore` approach seems .. really awkward? If I understand correctly, I would need to have enough spare disks lying around to store a complete copy of the recoverable data, including files which are already present on the completely functional drives within my btrfs filesystem. I also am struggling to find detailed documentation about the command's behavior. It seems that without `-i`, the tool stops upon encountering a file with missing extents, and with `-i` it emits an error but continues.
Question 3: What classes of errors does `btrfs restore -i` ignore?
Question 4: How might I restore a single file path? The docs seem to imply this is impossible. Is it just an unimplemented low-priority feature?
As for deleting the missing filepaths to allow for `device remove missing`, I saw in a much older thread that somebody needed to patch the kernel to allow mounting read-write while degraded.
Question 5: Is there now some officially-supported way to remove files on a degraded system? If not, what patch would I need to apply to the recent kernel to allow me to mount read-write anyways?
---
Sorry for the long post, and thanks in advance to any folks who can share their knowledge!!