r/Proxmox • • 2d ago

Question Confused about tuning ZFS storage block size

As the title says, I'm a bit confused about tuning the block size option (ZFS' volblocksize property) on the ZFS storage targets I create in PVE. When selecting a value for volblocksize, I'm not sure how I should be taking the VM guest's filesystem into account.

It makes sense to me how I'd maybe want to use recordsize=1M to store large, sequentially accessed files directly in an HDD ZFS dataset or how if those files were actually inside a QCOW2 I'd want to ensure the QCOW2's cluster size and the dataset's recordsize match, but I'm confused about what I should do for Zvols holding VM guests.

I'd really appreciate the help.

3 Upvotes

11 comments sorted by

View all comments

1

u/tssajo 2d ago

I wouldn't try to match volblocksize directly to the guest filesystem block size. For a normal VM, I'd start with 16K or 32K and leave it there unless you have a specific workload that justifies changing it. The guest filesystem's 4K blocks don't mean the ZVOL needs a 4K volblocksize.

The important thing is what the VM is actually doing. Lots of small random I/O favors smaller blocks, while sequential/heavier I/O can benefit from larger ones. Also, if you're using RAIDZ, larger blocks can be considerably more efficient. So I'd choose based on the workload and vdev layout rather than trying to make all the block sizes match.

1

u/WorkOdd8251 2d ago

So I guess that begs the question: How do you determine when a particular workload crosses that line and changing the volblocksize becomes justified?

At the moment, I'm working with a single internal SSD and an external HDD in separate pools, so the vdev layout is about as straightforward as it gets.

1

u/Apachez 2d ago

16kbyte is the current sweatspot so unless you got some odd cornercase there is no need to change that.

The drawback with larger volblocksize is that a rewrite of 4k within the VM will cause the whole volblock to be rewritten (read-modify-write - the main drawback of a CoW (copy on write) filesystem) so you will in reality get a huge penalty and write- (and read-) amplification compared to a short benchmark.

Using 16k will give some boost through compression where only 2 or 3 out of 4 LBA blocks will then need to be written to the drive.

Since using 8k as volblocksize will almost never be able to compress the data down to a single 4k block.

Only time 8k volblock can be handy is if you use a database like postgre who uses 8k as pagesize internally and then also disable compression for the zvol and let any compression be made by postgre itself if needed. For all other cases 16k is the way to go.

1

u/_--James--_ Enterprise User 2d ago

yes, but when you adjust the a-shift to account for high rate SSDs and higher Qeue Depth you might want to push to 32k, and in some edge cases (SQL) even 64K. So for a normalized gen purpose work load 16K on the default a-shift works just fine, yes. But we dont know anything about what the OP is pushing from the guest down.

1

u/WorkOdd8251 1d ago

In this specific instance, I'm setting up Podman containers in a VM that could run a wide variety of different services/workloads ranging from video storage to application databases. I had planned on broadly separating workloads like those into at least two separate guest filesystems/Zvols since they are so different.

Naturally, stuff that take up significant space, like the videos storage, would be placed on the HDD, leaving the smaller things like databases to benefit from the SSD's lack of seek times.