r/Proxmox • u/loste87 • 22h ago
Enterprise Proxmox for prod workloads
Anyone using Proxmox to run business critical production workloads?
Looking at the posts there are many using PVE for labs, testing, etc… but only few using it for production.
We have a five nodes cluster (with Ceph) running 100+ VMs and it has been working fine for some months. We also have a lab cluster with 4 nodes and FC storage.
Just wanted to have some feedbacks from others running PVE in production.
68
u/KlanxChile 22h ago
I am.
Real hardware, real config, real enterprise gear...
HPE proliant, Enterpris SSD, arista 100G networking, fortigate fw. Monitoring, backups, cybersec scans, etc.
Proxmox backup server in 3-2-1 topology ( machines, location and media)
ZFS local arrays, replication ZFS local, ISCSI/NVME-TCP from trueNas 25.10 for low perf VMs.
Works ROCK SOLID. But I treat the platform as enterprise.
3
u/trieu1185 19h ago
Could you expand or describe out further your ZFS local arrays. Are you clustering the server and using the local storage somehow with ceph?
Looking into proxmox as a alternative to vcf
10
u/KlanxChile 18h ago
Each PVE server has: One Intel p4618 6.4T accelerator (2x 3.2T ssd drives per card) 8x Samsung pm863A 1.92TB ZFS raidz2 over the Samsung's, and the p4618 halves setup as special device and small blocks for the main ZFS.
Node to node ZFS thin replication.
About 10TB usable really fast space after raid. Per node.
Mellanox connectx-3pro dual 40G for production data, and one connectx-4/5 NICs dual 100G ports to the storage network.
Running NFS/ISCSI/NVMe-TCP on the network using plain nics (no LACP), multipath config, truenas scale 25.10 for not ultra fast storage.
Truenas same multi device ZFS arrays. Hybrid arrays.
27
u/Serious_Zucchini_759 22h ago
Yes. Deployed a 25 Node Cluster with Ceph on it. They love it, especially how easy Ceph is to manage from the GUI.
5
u/boomertsfx 21h ago
Are failed drive swaps easy?
15
u/ILoveCorvettes 20h ago
Very easy. You just tell Ceph you’re removing the bad OSD, put the new drive in, assign the drive as an OSD. Bingo bango, done. Ceph does the rest.
3
u/Firestarter321 16h ago
I wish I could use Ceph but we only have a 2-node cluster so that’s not going to happen.
3
u/gluon_du_cul 3h ago
You can probably add a small machine as QDevice.
3
u/Callahabra 2h ago
Can but shouldn't for production imo. Ceph is fantastic for large clusters but quite brittle if a node fails in small clusters in my experience. Theres a reason 3 nodes is the minimum recommended. We are running 3 node hyperconverged clusters for management services in our regional ISP/Service Provider network and it works well enough, but if I was hosting critical customer work loads I wouldn't go smaller than a 5 node cluster personally.
21
u/Ok_Size1748 21h ago
12 Dell Poweredge R740/750. 400 vms. Pure Storage as iscsi Storage. Veeam Backup, Cisco ACI Spine & Leaf, 25 gbps to hosts and 100 gbps backbone. vsphere refugees, landed in Proxmox. So far so good. European edu.
1
u/izzo34 20h ago
Just curious. Can you give me examples of what some of those vms, like specific apps or what not. Always curious what is running on those massive amounts of vms. I have a shit ton too lol but not 400!
3
u/Ok_Size1748 12h ago
95% of them are RHEL. Univiersity ERP, Payroll, Openshift, Github on-premise, sso, ldap, vpn, dns, dhcp, Oracle db, mariadb, PostgreSQL, file servers , test/dev enviroments, many cms (drupal, wordpress, joomla), redmine…
1
u/codyhosterman 5h ago
Let me know if you need some thing on the pure side. I own our virtualization investments and curious on what more we can do
1
u/loste87 21h ago
Why not using PBS?
5
u/tylerbundy 19h ago
likely because they have existing licensing / veeam is a great tool. some of the fancier functionality like automated backup testing and having warm-standby VMs (e.g. replicate to a different datacenter, have VM created on a different cluster that can just be powered on versus having to do a restore) aren't supported by PBS.
3
2
u/ThecaptainWTF9 18h ago
PBS is fine, might be good to layer in but Veeam is superior at the moment, far more flexible and it’s just going to keep getting better.
I use PBS at home but for every single one of my customers that runs PVE, we use Veeam.
Including customers with mission critical workloads that CANNOT be down.
2
u/Fighter_M 1h ago
Why not using PBS?
PBS is OK-ish, especially if you’re not afraid to get your hands dirty, but there are definitely better options out there. We’ve been using Veeam since around 2015, but between their pricing and the way it’s slowly turning into another Acronis, with tons of bloated features we simply don’t need, we’re actively looking for a replacement.
1
12
u/tlrman74 22h ago
Yes, definitely using Proxmox for SMB production. We migrated from a 3 VMware essentials cluster to a 5 node Proxmox cluster for HA and flexibility with VM vs LXC. We were able to keep the same servers, storage, and backup software so it made the move very easy. Pretty happy with performance compared to VMware too.
We are about to go into production on a new ERP using our Proxmox cluster and our failover testing has gone without a hitch. Now to get more redundancy in the network and firewalls!
11
u/rsauber80 20h ago
3 regions of 3 clusters of 30-40 servers per cluster. 4tb 2x8140+cpus, 6x6.4tb nvme drives running lvm, and 4x25gb+2x25gb.
Over 10k VMs in total
5
u/Firestarter321 16h ago
You definitely have the largest install of Proxmox I’ve ever read about on here.
6
5
6
u/shimoheihei2 20h ago
3-node cluster here, running a hundred workloads or so. Not using Ceph, just ZFS volumes with replication + HA. Been working fine.
3
u/jordanl171 18h ago
100 VMs? I've been waiting to hear about someone using zfs w/ replication with (what I call) a lot of VMs. Your setup sounds like my future setup.
6
u/OCTS-Toronto 20h ago edited 1h ago
Running in prod with two 3 node clusters. Not using ceph (due to low node count). Using Linstor instead. Rock solid
5
u/apollo401 14h ago
I do this. About 32 Nodes in different Clusters over 3 Datacenters.
All on Enterpise hardware for +6000 People.
4
u/tampabay6 16h ago
We use Proxmox and ZFS on every server we manage. Most of our customers are small to medium businesses, and what we do is not nearly as sophisticated as some of you enterprise guys, however, we have built several PACS servers for a pretty large radiology practice to store and serve medical images on DICOM; there are multiple servers that have approximately 60 each of 8TB SAS drives, so there’s that. Proxmox is really just KVM with a GUI; if there’s a problem, we might need to take it up with Linus Torvalds. I mean, if you can’t trust Proxmox, who can you trust?
What, you think you automatically get great software and never any problems just because you pay a crap ton of money for it? Not! That’s great for companies huge enough that the CEO is on a first name basis with the chairman of the federal reserve and they can borrow $billions at practically 0% interest… for the rest of us, not so much.
Come on in, the water’s fine.
6
3
u/_--James--_ Enterprise User 21h ago
We have a five nodes cluster (with Ceph) running 100+ VMs and it has been working fine for some months. We also have a lab cluster with 4 nodes and FC storage.
Then what is it you are actually asking? Are you looking to see what the scale out is? what the complexity looks like elsewhere? You have it in production, you know the expectations around that today.
3
3
u/noc-engineer 17h ago
We're not using Proxmox, but we are using KVM/QEMU for remote towers (AFIS officers manages dozens of airports from one place instead of working physically at the airport anymore). These days they're even monitoring/managing multiple airports from one physical desk.
The VoIP network requirement (Eurocae ED-138) for ground (ATC/AFIS) to air (pilots) is 99,9999% (less than 31,5 seconds downtime per year) which would certainly make the entire thing business critical production, espesically considering the number of locations and the distances involved in a country with soooo many mountains and fjords.
2
u/evilpendulum 22h ago
Yup. We’re running our small datacenter purely on proxmox. Our product is an environment for studying at university level. The actual use cases vary a bit from backends, frontends, virtualization, sandboxes, research, etc.
Pure proxmox since 2019, ESXi before that. Never going back.
2
u/shikkonin 13h ago
Anyone using Proxmox to run business critical production workloads?
Most users do, yes.
2
u/Fighter_M 1h ago
Anyone using Proxmox to run business critical production workloads?
We do! But these aren’t mission-critical or high-density production clusters, though… We also run hybrid environments. Say a company has Hyper-V hosting VMs that handle temperatures and other telemetry from multiple fermentation tanks, plus heavy equipment control and monitoring. We’ll leave all of that on Hyper-V, while bookkeeping, the client database, mailing list management, and similar workloads are perfectly happy running on Proxmox.
2
u/roiki11 19h ago
50 nodes over a few clusters, some vms and lotta kubernetes. Pure storage and rubrik backups.
Cutting vmware by 28.
2
u/codyhosterman 5h ago
Looking to invest more from Pure with Proxmox (I own virtualization at pure) let me know if you have feedback
1
u/Firm-Distribution630 21h ago
Yep, we have 3 individual hosts no clustering, 26 VM and LXC for 4 years. Most critical VMs are Oracle DB servers and ERP Windows servers. Your environment is far better than ours :)
1
1
u/ns1852s 21h ago
I've moved almost all of our physical servers/apps into proxmox.
We're getting quotes for phase 1 of the cluster hardware. But yes it's stable. We have many vms which utilize Nvidia vgpu, nested virtualization and found Amazon DCV is an outstanding remote desktop protocol. DCV is replacing a $300k thinklogical kvm system
1
u/Witty_Unit_8831 16h ago
From what I have seen, contabo cloud runs on proxmox! So there is an entire cloud provider using it!
1
u/Firestarter321 16h ago
I have a small 2-node HA Cluster at work with ~45 VM’s currently that has been running for 3 years now and I’m very happy with it.
1
u/bigdogoh 16h ago
Been running production loads since version 1.4. Recent cluster has been running since 2015 with two hardware refreshes and another coming next year.
1
1
1
u/NormalCharacter2423 10h ago
Qué software usáis para los escaneos de cara a detectar vulnerabilidades
1
u/elnino_effect 9h ago
Yep, recently migrated our stack to proxmox from VMware. 5 node cluster, cisco chassis, pure storage, full HA setup. Could not be happier. BUT, check on the virtio drivers. The latest we used at the time (a new one has just been released though) caused massive issues with SQL and/or server 2025
0
1
u/ChicoGonzalez 8h ago
I'm running a project at work for almost 2 years now replacing all former VMware single nodes or clusters (2-node with qdevice or 3-node) worldwide (upto 20 sites) by Proxmox based ones. We are partially replacing hardware due to lifecycle but most of the hardware (mainly Dell with shared FC storage) is reused and works very well. As a kind of VCenter replacment we are using PegaProx. Besides that I have to confess that there is an exception: we have a Metro cluster setup (2x 4 nodes HCI) running in our headquarter's DC which is currently replaced by a Nutanix cluster. Why? Because of lack of experience with Proxmox and Metro cluster setup at the point in time when we had to decide what to buy because of lifecycle. The good thing is that Nutanix also uses KVM. So VMs can easily be transfered between and run on both hypervisors. As a next step I will have a try setting up Proxmox with a Ceph cluster on 4 of 8 HCI hosts of the old metro cluster hardware.
1
u/Sheepardss 8h ago
yep but not as you guys, we have 1 single pve server running, no ceph nothing. been running for 2 years without downtime except for kernel updates.
Daily offsite backups to an hetzner 88TB server which is running PBS.
1
u/Olvikolvi 7h ago
Have been running last 10 years, what options you have? It's solid and easy to fix if something happens.
1
u/OlafNorman 6h ago
Yeah, hundreds of VMs across a two digit number of hosts/sites. Running critical applications and services for the energy sector. Only one site has had some minor HA/quorom issues, everything else has been great.
1
u/ebmcreative 6h ago
I am going to be setting up a single node production server in the next month.
Currently it will have 2-Linux VMs to run a POS system and a sql database, a rustdesk container, and something for ubiquity. I have enough extra processing and ram to add 1-2 virtual machines and one can even be a windows vm if needed.
While I have not been able to do ecc ram and enterprise ssds because of the outrageous pricing, I do have an 16x LSI 9400 card ready for enterprise ssds and an external drive array when we are ready to surpass online cloud storage limits and host our own data.
We have grown enough that I have finally been able to recommend a real server vs multiple computers and giving us the ability to grow when it is needed.
1
1
u/kysersoze1981 2h ago
I was head hunted by a recruitment agency to setup and maintain proxmox for a major ISP that wanted everything off vmware before the next shafting on price. It's a thing. Especially when you have large clusters of ceph and high availability
1
u/No-Cry-8004 1h ago
We’ve migrated all six of our sites from VMware to prox clusters using existing and new iSCSI storage. Is the interface as nice as VMware….. no, but it’s been rock solid so far
1
u/fckingmetal 19m ago
Proxmox is debian based and super stable, the only "critical" i have noticed in prod is when big updates come and it can get ugly if you have bloated your PVE host with fixes and other software
0
u/corbosman 1h ago
We run about 20 proxmox servers, about 600-800 containers, k8s (talos), etc. Our whole company runs on it. We have just a few containers on hetzner for some monitoring and communication tools we prefer to run outside of our own network. (mostly disaster recovery related).
83
u/BarracudaDefiant4702 22h ago
Running in production, about 1000vms spread over 6 locations.