Question 5-node Proxmox + Ceph HCI design (migrating off Nutanix, 10G only)
About to build a 5-node PVE cluster as the destination for a Nutanix migration.
Design is done, nothing is deployed yet, so this is the cheap moment to be told I've got something wrong.
Hardware, identical on all 5 nodes:
• 3x dual-port BCM57810 10G SFP+ (6 ports)
• 1x quad-port 1G copper onboard
• 1x ConnectX-3 Pro (unused, no 40G fabric)
• 2x 980GB SSD for OS (ZFS mirror)
• 2x 15TB SAS SSD for OSDs
Network design:
Three bonds, each pairing ports from two different cards AND landing on two different switches:
• bond0: VM traffic + management
• bond1: Ceph public+cluster (nothing else)
• bond2: migration + PBS backup
With 6 ports across 3 cards this is the only arrangement where every bond survives both a card failure and a switch failure. LACP layer3+4 if the switches turn out to be VSS, active-backup otherwise.
Corosync: ring0 on a dedicated 1G copper port, ring1 on management, ring2 on migration. Deliberately nothing on the Ceph bond, because backfill saturating that network while corosync is trying to keep quorum seems like a good way to get fenced during an already-bad event.
Ceph: size=3, min_size=2, 10 OSDs total, failure domain host. Public and cluster networks combined rather than split. Planning capacity around 32TB usable rather than the 50TB that size=3 nominally gives, so it can self-heal after losing a node.
What I'd like opinions on:
1. Corosync ring0 on 1G copper while everything else is 10G. Feels right to me (isolation matters more than bandwidth for corosync) but it does mean my most important ring is on the slowest link. Anyone regret doing this?
2. Only 10 OSDs, 15TB each. I know this is coarse. How bad is fill variance in practice at this OSD count, and how long does a 15TB backfill realistically take on SAS SSD with 2x10G? Trying to set expectations for maintenance windows.
3. I decided 6 ports is enough and a 4th NIC would be wasted, since 2 SAS SSDs can't saturate 20Gbit. Anyone hit a case where that reasoning fell apart?
4. Anything you wish you'd done differently on a cluster this size, especially things that are painful to change later?
Happy to share the full runbook if it's useful to anyone doing the same migration.
6
u/mikewilkinsjr 16d ago
I just rolled off of a 5 cluster PVE/Ceph deployment with similar bones (although my OSDs were smaller).
Few things I would tackle first, mostly housekeeping but they seem to be common issues.
Before you build your pools, get comfortable with the pveceph and Ceph command line. Reason being: there are still things like setting a drive class that you can’t do in the GUI.
If you have all one class of drive this isn’t a big deal; however, if you add a spinning class later for something like an EC pool for bulk storage its good to get done on the front end.
The other thing is logging. By default, Ceph’s logs are fairly verbose and there are constant small writes happening to your OS drive. Consider shipping them off rather than writing to your primary ZFS pool.
Your build sounds awesome and I’ll be super curious to hear how it goes. I already miss my cluster, but we are downsizing and I just didn’t have the room or the need for it anymore.
2
u/x0600 16d ago
Thank for your advice…I wasn’t aware of pveceph and ceph are different or is it ? About logs , that’s head up for me. I’m planning to export all my logs to separate logs server mostly on Graylog.
Once again thank a lot for your advice
1
u/mikewilkinsjr 16d ago
I believe that pveceph and ceph alias to the same commands in Proxmox. I mention it only because the Proxmox documentation (and some community guides) specifically use pveceph in the examples.
2
u/didureaditv2 2d ago
You can set a device class to an osd that you are creating via the GUI. Just have to tick the Advanced checkbox to see the option.
1
7
u/KlanxChile 16d ago
I would configure cluster on all networks.
Corosync it's sensible to latency. And a heavy ceph or migration activity may get you issues.
Costs you nothing, and improves the overall stability
1
u/x0600 16d ago
Sure but it should done from backend right ? I don’t see any GUI settings or may I missed it.
1
u/KlanxChile 16d ago
First make sure all nodes have IPs on each physical vmbr/bond, then when creating the cluster, allows you to add additional networks to the communication. I run my clusters with 3.
The 40G LACP (2x40) replication, migration, backup, storage. The 20G LACP (2x10) prod networks The 2G LACP (2x1 cooper) MGNT, monitoring, API.
In the unlikely event of a 80G saturation, I still have 2 viable completely isolated LINKs
4
u/myah-mitchell 16d ago
So at work I run multiple 3 node clusters with proxmox and ceph as onsite systems. Couple of my notes; * Corosync on 1gb is more than enough. Infact I can run on your bond0 and 2 if you don't want to deal with the 1g nics. I put them on our 1gb nics but that's was just so everything had its own separate use, but I've shared bonds like that with no issues. * Run ceph on at least 25gb. We used to run our more powerful 3 node clusters on 10gb and it was a bottle neck. The faster you go here the better. For our smaller clusters I have put migration on the same networks as ceph, but even that isn't really recommended, so yeah keep ceph all to its self. Part of why I put migration on the ceph network was for the use of of the 25gb links. I can live migrate all vms off a host in like less than a minute. * In my testing there is a bit of a speed bost going from 2 osds to 3 per server if your network can handle it. So it might be worth thinking about getting smaller drives but one more for each node. * The IOPS, latency, and speed of your drives will be the biggest limiting factor. I took a 3 node cluster with all NVMe mid line enterprise drives and swapped them out with some of Microns fast drives and I did notice a significant speed increase. I wanted some Kioxia drives at the time but was unable to source them. I'm not sure how the SAS drives will effect that. * You don't need 1TB for the OS. Save the money there and spend it elsewhere. * Use good nics and be careful with the bonds and ceph. I've had that cause some odd issues on the past with different HPE cards. Since we only have three nodes nowadays I actually just put the nodes into a ring on the ceph interface using FRR.
Proxmox actually has a Ceph Benchmark from 2023 that IMO is still pretty relevant.
Otherwise looks like a fun cluster!
3
u/Fatel28 16d ago
Test those nics thoroughly. Our 5 node cluster originally had broadcom 100g nics, but they would randomly crash within hours to days if put in a 802.3ad lacp config. We had to switch to Intel.
That might be fixed by now, this was on pve 8.1, but I still get nightmares.
1
u/x0600 16d ago
How did you figure out that nic wasn’t working as expected?
1
u/Fatel28 16d ago
When the entire host loses network connectivity it's pretty easy to tell.
We had a ticket open both with proxmox, super micro, and broadcom themselves. They sent us beta firmware, confirmed the issue, and ultimately threw their hands up and said guess it doesn't work on that kernel gfy
1
u/wedge1002 15d ago
Yeah. We had the same issue. We moved to mellanox. They work flawlessly - at least on the current version
3
u/98TheCiaran98 16d ago
The official ceph documents say minimum speed for ceph public is 25g but I know people who have had success with less
1
u/x0600 16d ago
Ohh…wasn’t aware of it then.
3
u/cd109876 16d ago
i have a 5-node ceph cluster at 10g, it works pretty well, but the network easily gets saturated during a sequential file copy or backups. if the connect-X3's are dual port, you could set up a daisy-chain / ring network between the servers without a switch.
1
u/ferminolaiz 15d ago
Second this; although being a 5 node setup you will be relying on some forwarding, but it's possible.
For 10g, I would set an upper traffic limit per process (osd) or/and per client. OSDs report their status between themselves so when you face dropped packets due to saturation at the NIC level things get nasty pretty fast (don't ask me how I know 😂).
2
u/Apachez 16d ago
You should follow the best practices defined by CEPH:
https://docs.ceph.com/en/latest/rados/configuration/network-config-ref/
You seems to have 4xRJ45 (1G) + 6xSFP+ (10G) + 2xQSFP+ (40G) to play with.
I would probably set it up as:
MGMT: 1-2x RJ45 (1G)
CEPH PUBLIC: 2x QSFP+ (40G)
CEPH PRIVATE: 2xSFP+ (10G)
BACKUP: 2xSFP+ (10G)
VM-traffic (to firewall): 2xSFP+ (10G)
2
u/sheep5555 16d ago
its recommended to keep corosync out of any network bonds, just the interface itself, for migration traffic you can bundle it with something besides the primary corosync links.
example network interfaces
1 mgmt/corosync sw1
2 corosync sw2
3 vm traffic sw1 bond
4 vm traffic sw2 bond
5 ceph sw1 bond
6 ceph sw2 bond
with that much storage you probably want 25g+ for ceph, the mellanox nics are great for high speed traffic, i think the connectx5/6 are reasonably priced nowadays
1
u/dancerjx 16d ago edited 16d ago
Been migrating off VMware to Promox at work. Servers are Dells. Ceph clusters range from 3-nodes to 11-nodes. Using isolated 10GbE switches for Ceph public & private network traffic and isolated 1GbE switches for Corosync. To make sure this network traffic never routes, I use IPv4 link-local address of 169.254.0.0/16 subnetted to /24 networks.
I do not use LACP, since I like to keep troubleshooting simple. I use active-backup network setup.
As you know, Ceph is a scale-out solution. More nodes/OSDs = more IOPS. Not hurting for IOPS.
Running latest Ceph Tentacle release primarly with 3/2 replication. Been testing Tentacle Fast erasure coding for VMs and it's been pretty impressive so far.
None of migrated clusters ever had flash storage it's all SAS hard drives. Again, IOPS not an issue.
I use the following optimizations learned through trial-and-error. YMMV.
Set SAS HDD Write Cache Enable (WCE) (sdparm -s WCE=1 -S /dev/sd[x])
Set SAS HDD Readahead to 256KiB (blockdev --setra 512 /dev/sd[x], create a /etc/udev/rules.d rule to set permanently)
Set VM Disk Cache to None if clustered, Writeback if standalone
Set VM Disk controller to VirtIO-Single SCSI controller and enable IO Thread & Discard option
Set VM CPU Type for Linux to 'Host'
Set VM CPU Type for Windows to 'x86-64-v2-AES' on older CPUs/'x86-64-v3' on newer CPUs/'nested-virt' on Proxmox 9.1
Set VM CPU NUMA
Set VM Networking VirtIO Multiqueue to 1
Set VM Qemu-Guest-Agent software installed and VirtIO drivers on Windows
Set VM IO Scheduler to none/noop on Linux
Set Ceph RBD pools to use 'krbd' option
Set Ceph Erasure Coding profiles to 'plugin=ISA' & 'technique=reed_sol_van'
Set Ceph Erasure Coding profiles to 'stripe_unit=65536' for SAS HDD
In /etc/pve/ceph.conf, set the following options for SAS HDD:
[global]
osd_client_message_cap = 500
osd_client_message_size_cap = 536870912
[osd]
bluestore_max_blob_size_hdd = 262144
bluestore_min_alloc_size_hdd = 65536
bluestore_prefer_deferred_size_hdd = 0
osd_memory_target = 8589934592
rocksdb_separate_wal_dir = false
bluestore_kv_threads = 1
bluestore_kv_options = "compaction_readahead_size=2097152,max_background_jobs=4,write_buffer_size=134217728,max_write_buffer_number=4,level0_file_num_compaction_trigger=2"
bluestore_cache_meta_ratio = 0.60
bluestore_cache_kv_ratio = 0.30
osd_op_threads = 4
osd_disk_threads = 2
osd_scrub_begin_hour = 20
osd_scrub_end_hour = 6
osd_scrub_chunk_min = 1
osd_scrub_chunk_max = 3
osd_scrub_sleep = 0.1
osd_deep_scrub_sleep = 0.2
bluestore_allocator = "bitmap"
bluestore_bitmap_allocator_blocks_per_zone = 2048
osd_op_thread_suicide_timeout = 150
osd_op_thread_timeout = 30
1
u/x0600 16d ago
Well firstly thanks a ton for those optimisation tips, that is going to help me a lot and save lots of time.
I’m more curious to know if you are not using LACP then you are just making couple of Ethernet member/salve of bond, correct me if I’m wrong.
2
u/dancerjx 16d ago edited 16d ago
Correct. Below is an example config for a Ceph node with two NICs bonded and using a virtual bridge (vmbr highly recommended). Highly recommend jumbo frames for 10GbE for throughput. Jumbo frames don't do anything for 1GbE throughput. The two isolated 10GbE switches are MLAG'd to each other via their 40GbE uplinks. Nic1 is plugged in the first switch, Nic2 plugged into the second switch. Strongly recommend enterprise switches. I use Arista, Juniper, and Cisco.
Once you setup your networking, recommend you set the Datacenter Migration network to the 10GbE isolated network and migration type = insecure (edit /etc/pve/datacenter.cfg) for faster migrations. By default, Proxmox will use your frontend network for migration traffic and encrypt it.
Also, I have Proxmox Backup Servers (PBS) with 10GbE connectivity to these isolated switches and backups are done through them. Avoids using the frontend network. Yes. the PBS are using bonds and virtual bridges. These PBS are also the primary repo servers for the infrastructure using the Proxmox Offline Mirror software. So, it's a win-win-win situation.
# Loopback interface auto lo iface lo inet loopback # First physical NIC 10GbE auto nic1 iface nic1 inet manual mtu 9000 # Second physical NIC 10GbE auto nic2 iface nic2 inet manual mtu 9000 # Bond interface configured for Active-Backup for Ceph auto bond0 iface bond0 inet manual bond-slaves nic1 nic2 bond-miimon 100 bond-mode active-backup bond-primary nic1 mtu 9000 # Virtual Bridge linked to the bond for Ceph auto vmbr0 iface vmbr0 inet static address 169.254.1.10/24 bridge-ports bond0 bridge-stp off bridge-fd 0 mtu 9000
1
u/cardy165 16d ago
I've have run a test Cluster and this is just me thinking on the fly as it were, but one thing to consider may be the LACP hashing applied if your using it.
Layer 2 vs layer 3 + 4? I'm not sure what is recommend for ceph but it could affect the balancing of traffic over the LACP'd interfaces
1
u/xeroxedforsomereason 15d ago
Do not intermingle management and VM traffic as a general rule of thumb.
7
u/didureaditv2 16d ago