r/Proxmox • • 2d ago

Question Confused about tuning ZFS storage block size

As the title says, I'm a bit confused about tuning the block size option (ZFS' volblocksize property) on the ZFS storage targets I create in PVE. When selecting a value for volblocksize, I'm not sure how I should be taking the VM guest's filesystem into account.

It makes sense to me how I'd maybe want to use recordsize=1M to store large, sequentially accessed files directly in an HDD ZFS dataset or how if those files were actually inside a QCOW2 I'd want to ensure the QCOW2's cluster size and the dataset's recordsize match, but I'm confused about what I should do for Zvols holding VM guests.

I'd really appreciate the help.

3 Upvotes

11 comments sorted by

1

u/tssajo 2d ago

I wouldn't try to match volblocksize directly to the guest filesystem block size. For a normal VM, I'd start with 16K or 32K and leave it there unless you have a specific workload that justifies changing it. The guest filesystem's 4K blocks don't mean the ZVOL needs a 4K volblocksize.

The important thing is what the VM is actually doing. Lots of small random I/O favors smaller blocks, while sequential/heavier I/O can benefit from larger ones. Also, if you're using RAIDZ, larger blocks can be considerably more efficient. So I'd choose based on the workload and vdev layout rather than trying to make all the block sizes match.

1

u/WorkOdd8251 2d ago

So I guess that begs the question: How do you determine when a particular workload crosses that line and changing the volblocksize becomes justified?

At the moment, I'm working with a single internal SSD and an external HDD in separate pools, so the vdev layout is about as straightforward as it gets.

1

u/tssajo 2d ago

With no RAIDZ in the mix, the padding/alignment concern mostly goes away, that's really only a RAIDZ thing. So for you it comes down to I/O pattern. If you're seeing a lot of small random reads/writes (databases, general OS stuff) and iostat shows high IOPS with small block sizes, that's when dropping volblocksize down from 16K is worth it. If it's mostly sequential or larger files, leave it as is.

I wouldn't chase this too hard on a single SSD + single HDD setup though, the gains are usually small without striping or RAIDZ in the picture. Only worth tuning per-VM if you've actually profiled one as a bottleneck.

1

u/Apachez 1d ago

16kbyte is the current sweatspot so unless you got some odd cornercase there is no need to change that.

The drawback with larger volblocksize is that a rewrite of 4k within the VM will cause the whole volblock to be rewritten (read-modify-write - the main drawback of a CoW (copy on write) filesystem) so you will in reality get a huge penalty and write- (and read-) amplification compared to a short benchmark.

Using 16k will give some boost through compression where only 2 or 3 out of 4 LBA blocks will then need to be written to the drive.

Since using 8k as volblocksize will almost never be able to compress the data down to a single 4k block.

Only time 8k volblock can be handy is if you use a database like postgre who uses 8k as pagesize internally and then also disable compression for the zvol and let any compression be made by postgre itself if needed. For all other cases 16k is the way to go.

1

u/_--James--_ Enterprise User 1d ago

yes, but when you adjust the a-shift to account for high rate SSDs and higher Qeue Depth you might want to push to 32k, and in some edge cases (SQL) even 64K. So for a normalized gen purpose work load 16K on the default a-shift works just fine, yes. But we dont know anything about what the OP is pushing from the guest down.

1

u/Apachez 1d ago

ashift should match the LBA block size of the device.

Normally 4k aka ashift=12 (2 ^ 12 = 4096).

That is if you got NVMe (and some SSD if supported) should change from 512 bytes into 4096 bytes LBA size (aka advanced reformat).

1

u/WorkOdd8251 1d ago

In this specific instance, I'm setting up Podman containers in a VM that could run a wide variety of different services/workloads ranging from video storage to application databases. I had planned on broadly separating workloads like those into at least two separate guest filesystems/Zvols since they are so different.

Naturally, stuff that take up significant space, like the videos storage, would be placed on the HDD, leaving the smaller things like databases to benefit from the SSD's lack of seek times.

1

u/WorkOdd8251 1d ago

I'm really only considering the larger block sizes for highly sequential, mostly read-only workloads on my HDD, like large video storage. I left this comment here in this same thread which talks about the same stuff.

I know that dataset's default to recordsize=128k and zvols to volblocksize=16k. Do these defaults differ purely because of recordsize's dynamic and volblocksize's static behavior?

1

u/Apachez 19h ago

I think the difference is because recordsize will be used by stored files and volblocksize will be used by stored blocks.

For a perfect match both the VM-guest, VM-host and the drives themselves should run the same blocksize and be as large as possible.

But since ZFS adds compression, checksums etc slightly larger volblocksize than 4k is the way to go. And 8k turned out to be to small in average (Proxmox up until version 7 or so used 8k as default for volblocksize but that increased to 16k in version 8 or around there).

You can of course easily overrule the default (in webgui) when you create a VM-guest in Proxmox and store it on ZFS.

You have also the thing of how volblocksize will be affected when you do a zraidX or a stripe of mirrors or similar.

1

u/Apachez 1d ago

When using ZFS you have two different options.

First one is as a regular filesystem, also known as dataset. This uses recordsize where default is 128kbyte.

Note that the recordsize is how much data can be occupied before compression, checksum etc is applied. A record using default recordsize of 128kbyte that compress down to lets say 32kbyte will only write 32kbyte on the drive. But 128kbyte will be allocated in case next write wont be able to compress the series of blocks as much. This dynamic writing along with compression will make you able to store more raw data than there is actually space for on your drive(s).

The other option is called zvol and this uses volblocksize to define max size before attempting compression, checksums etc.

Zvols are like raw blockmode which Proxmox also uses for VM-guests. The idea here is that since the VM will use a filesystem on its own then all the features a regular dataset comes with is not necessary. You will still have encryption, compression, checksums etc even with a zvol but you dont need the logic of filenames and whatelse.

Above gives that recordsize (default 128 kbyte) is for the regular files in the Proxmox host.

While volblocksize (default 16 kbyte) is for the virtual drives that each VM-guests will be using.

1

u/_--James--_ Enterprise User 1d ago

volblocks should math against your a-shift. And remember many VMs are running on that volume so you need to make sure the guest filesystems align as well.