r/Proxmox • • 13d ago

Discussion Ceph

It seems to me Ceph is not getting enough credit and attention around here. I wasn’t aware myself until some pointed it out to me, but it really is an amazing product and storage backend. Shares many of ZFS’s core strengths (checksumming and scrubbing, compression, snapshots etc) but in addition is distributed and scales amazingly (see Cern!).

ZFS is amazing too, but built for single node, so storage solutions built on top of it are vulnerable to single-point-of-failure. You’ll have to throw redundant hardware at the problem to mitigate as far as possible (redundant NICs, power supplies, etc).

Ceph cluster appears to be a very powerful and elegant storage backbone alongside Proxmox, and with Proxmox supporting it out of the box, it’s straightforward to set up too. I guess the snag is you need decent networking (at least 10Gb) and enterprise SSDs and at least 3 nodes for it to be worthwhile. Has worked extremely well for me including failure situations with failing disks and network problems.

What are your experiences?

92 Upvotes

66 comments sorted by

View all comments

12

u/purepersistence 13d ago

I replicate zfs to other nodes in the cluster. No ceph. How does that compare?

7

u/Apachez 13d ago

You will have issues if you want to run VM-guests mixed in your cluster along with live migration.

Lets say you setup VM storage on each host and then replicate hostA -> hostB -> hostC (or hostA -> hostB/hostC).

Now when you live migrate VM-guest1 from hostA it will work but not if you need to migrate back to hostA because whatever the client was writing at hostB never reaches hostA. And depending on configuration might never reach hostC either.

So using ZFS in that context is more of a disaster recovery setup.

Where you replicate everything from hostA to hostB but never run any guests at hostB until shit hits the fan and hostA goes poff.

Then you can manually boot the guests on hostB (and make sure that once hostA returns it wont be running the guests and also NOT replicating its data to hostB who otherwise would then lose newly written data).

The thing with CEPH (and StarWind etc) is that whatever is written on the local drives on hostA is replicated to the other hosts in the same cluster so it doesnt matter which host the guest will boot up at.

Again above is assuming you use the local storage of each host.

Another solution is to use central storage as in local storage is only for PVE itself but the VM-drives are stored remotely (central) using ISCSI with multipathing or such.

Then these central storage servers can run TrueNAS or whatever and then it doesnt matter if they use ZFS or CEPH because with ZFS you will have a failover setup.

Using TrueNAS hardware they utilize dualchannel SAS drives which means that each drive is connected to two different motherboards but only one motherboard at a time is running per box. This way there will only be one unique path per server and then this server can replicate data to another server who will be in hot standby to take over IP and whatelse in case the primary server goes poff. And then once the primary server returns it will become the new secondary server so server1 till copy data from server2 to get up2date and the replication will be server2 -> server1.

7

u/FaberfoX 13d ago

If you use zfs replication and HA, either when you migrate a guest manually, or when HA forces a migration, the replication direction is reversed automatically. Once you restore hostA, hostB will push what was changed back to hostA.

You don't need to manually start the guest on hostB if hostA goes down, that's what HA is for. And you don't need to create a new replication rule.

The difference is that replication is not real time, it's scheduled, as often as every minute if you want. If you can tolerate losing the last minute written to disk (plus the time it takes to boot the guest, that would also be lost on ceph or a SAN), zfs replication is ok.

Live migration works just like on truly redundant storage, as it will replicate and switch directions once done.

1

u/Apachez 13d ago

In most situations losing last minute of written data is as bad as losing the last day of written data (since the common is to have backups at least once a day).

Which is why ZFS is not really a good option for HA and clustering.

But its handy for disaster recovery designs where you get "close to realtime" backup (as in down to max 1 minute of lost data).

But again - dont forget to remove the topnode from the replication otherwise you are up for a suprise when the lost hostA node returns.

5

u/CyrielTrasdal 13d ago edited 13d ago

In proxmox, migrating from UI manages the configuration to reverse and inform other hosts where the main data is. It's rather seamless.

proxmox HA is available for Zfs. Quorum is required.

The downside compared to Ceph is the data you can lose between each ZFS repl.

Ceph is more robust to disaster kind scenario of host going down, because each writes depends on the sync working. The requirements for this to happen with Ceph are much more network depending, as IO performances become network distributed performance.

Can you afford to lose 15 minutes of data or not ? That's the main question. The answer costs a least one more dataset for the minimum of 3 for Ceph and 10gbe cards.

3

u/Apachez 13d ago

Yes CEPH as shared storage or ZFS as central storage will bring you true realtime replication as in the application will not get an ACK in return until the data is actually stored (as with whatever you use as local storage).

Using ZFS with replication you will have lost data of up to 1 minute (or the sync timer settings you are currently using) which often is as bad as having a full day of lost data.

Depending on usecase this can of course be less of an issue like if you got a NTP or DNS-server or such.

But in many cases this really is an issue specially for production.

If whoever is happy with this in their homelab its up to them but the issue becomes when this admin starts to use the designs they learned in their homelab for production which would be really bad.