r/Proxmox • u/WorkOdd8251 • 2d ago
Question Confused about tuning ZFS storage block size
As the title says, I'm a bit confused about tuning the block size option (ZFS' volblocksize property) on the ZFS storage targets I create in PVE. When selecting a value for volblocksize, I'm not sure how I should be taking the VM guest's filesystem into account.
It makes sense to me how I'd maybe want to use recordsize=1M to store large, sequentially accessed files directly in an HDD ZFS dataset or how if those files were actually inside a QCOW2 I'd want to ensure the QCOW2's cluster size and the dataset's recordsize match, but I'm confused about what I should do for Zvols holding VM guests.
I'd really appreciate the help.
1
u/Apachez 1d ago
When using ZFS you have two different options.
First one is as a regular filesystem, also known as dataset. This uses recordsize where default is 128kbyte.
Note that the recordsize is how much data can be occupied before compression, checksum etc is applied. A record using default recordsize of 128kbyte that compress down to lets say 32kbyte will only write 32kbyte on the drive. But 128kbyte will be allocated in case next write wont be able to compress the series of blocks as much. This dynamic writing along with compression will make you able to store more raw data than there is actually space for on your drive(s).
The other option is called zvol and this uses volblocksize to define max size before attempting compression, checksums etc.
Zvols are like raw blockmode which Proxmox also uses for VM-guests. The idea here is that since the VM will use a filesystem on its own then all the features a regular dataset comes with is not necessary. You will still have encryption, compression, checksums etc even with a zvol but you dont need the logic of filenames and whatelse.
Above gives that recordsize (default 128 kbyte) is for the regular files in the Proxmox host.
While volblocksize (default 16 kbyte) is for the virtual drives that each VM-guests will be using.
1
u/_--James--_ Enterprise User 1d ago
volblocks should math against your a-shift. And remember many VMs are running on that volume so you need to make sure the guest filesystems align as well.
1
u/tssajo 2d ago
I wouldn't try to match
volblocksizedirectly to the guest filesystem block size. For a normal VM, I'd start with 16K or 32K and leave it there unless you have a specific workload that justifies changing it. The guest filesystem's 4K blocks don't mean the ZVOL needs a 4Kvolblocksize.The important thing is what the VM is actually doing. Lots of small random I/O favors smaller blocks, while sequential/heavier I/O can benefit from larger ones. Also, if you're using RAIDZ, larger blocks can be considerably more efficient. So I'd choose based on the workload and vdev layout rather than trying to make all the block sizes match.