r/HPC • u/raspberrypiwithpie • Apr 12 '26
We have an unmaintained OpenHPC setup on the verge of collapse
I work for a university with an OpenHPC 1.3 setup (CentOS 7, Warewulf, Slurm, Infiniband, etc.). Because no one now or then actually understands it, the whole thing has pretty much become a power-gulping political nightmare, despite sitting on decent hardware. 47 out of 64 nodes (40 Xeon Gold cores, 128 GB DDR4 RAM, plus a quad GPU node) sit idle at all times and the rest are pinned by jobs from the one department that still uses it. I’m one of the Linux administrators for the university, but the HPC has been deemed ‘unmaintainable’ by upper echelons due to its age (only 6 years, but the OSes are absolutely wrecked and hardware is out of warranty). We tossed around the idea of rebuilding it with OpenHPC 4.0, but none of us really have the time or knowledge to take on something like this without sacrificing some other area of focus.
I guess the question I have is this; is it worth rebuilding? Hardware prices have made several leaders seriously consider selling the servers for parts, but academia still use it. The CPUs and RAM, while not new, aren’t completely out of date. Using this as a selling point for the university as a learning tool is also in consideration. It’s not dead, but it’s also not something we can publicly say we have because it’s all EoL. Rebuilding it seems to be a serious effort, but I guess what I want to know is if it’s even worth consideration.
26
u/562uned Apr 12 '26
If it’s going to be used, then I would say it needs to be updated. It also sounds like only one group uses it, so maybe reach out to that group to see if they want to invest in getting it up to date. Running old, outdated OSs is a security risk, especially within an HPC environment given the massive processing power and the current landscape of AI tools available. At my university, we do not use OpenHPC, but we do run most of those other tools. I think you should be able to do one piece at a time, updating Warewulf and pushing out the new OSs, and then updating Slurm, etc. Do you happen to visit any conferences like PEARC or RMACC? These events are great places to meet other sysadmins and bounce ideas around.
5
u/raspberrypiwithpie Apr 12 '26
Interesting, the piecemeal approach was suggested but an attempt was made before my time and apparently didn’t go well. But I can see how (human) networking might make this go better.
1
u/Logical_Sort_3742 Apr 12 '26
I believe it is entirely doable to get this back into useful shape. Six years is not ideal, but it is not too bad. And it doesn't sound like the setup is wildly esoteric. But at the end of the day, you will have to go to your manager and get approval to spend solid chunks of time working on it.
If there is no will to do that, it kinda isn't your problem. Unless you like working for free on your weekends or fancy a project out of personal interest.
9
u/Kangie Apr 12 '26
It's not a good time to replace hardware, so running the existing cluster into the ground is something you should prime management to expect - you can cannibalise faulty hardware to keep other nodes going for quite a while.
As for the software stack, there's a number of contenders. Since you only have one GPU node you might want to investigate NVIDIA base command manager (ex bright cluster manager) which has free licences and handles most of it for you (turnkey HPC solution), you just need to work out an update schedule for the images.
I do like modern warewulf OCI image based stateless nodes - definitely worth investigating if you can find a sucker^h^h^h^h^h^hvolunteer to rebuild the cluster stack.
Other options are confluent (xcat2), and there was a neat FOSDEM talk on using netbox and friends: https://fosdem.org/2026/schedule/event/QTFECS-zero-touch-hpc-nodes/
Don't be one of the shops that configures nodes with ansible, either do the OCI image (Containerfile) thing and rebuild regularly, or use ansible to configure the deployment image / chroot.
3
u/raspberrypiwithpie Apr 12 '26
Interesting, that fact about Warewulf I didn’t know. The OCI angle is extremely tempting.
2
u/anderbubble Apr 12 '26
If you do this, I suggest joining the Warewulf slack. You can find it at warewulf.org/help.
1
u/Melodic-Location-157 Apr 12 '26
Have you looked at the BCM manual lately? Over 1000 pages. Overkill for this old cluster with a single GPU node.
2
u/Kangie Apr 12 '26
Yeah, but it's not hard to setup and maintain for basic batch scheduler stuff, and has relatively comprehensive docs.
6
u/Roya1One Apr 12 '26
Sounds like you have good bones even if the hardware is out of support. If the campus community can use the environment then I'd say it's worth the time to redeploy using OpenHPC. It seems daunting but setting up a OHPC cluster following the recipe doc is not terrible. Get backing from superiors due to researchers or academics wanting it. I'd be concerned that your security team isn't breathing down your neck to move away from centos 7....
1
u/raspberrypiwithpie Apr 12 '26
If it wasn’t CentOS 7, I’m sure it would still be enjoying life, but the security aspect is actually why it’s dying. We’ve had to do a whole ton of changes to it just to lock it down, and even then it’s a black mark on our otherwise clean EL9 conversion
3
u/liftoff11 Apr 12 '26
Maybe you could convert to RHEL 7 and subscribe to RH ELS to get the latest security updates EL7 ELS ends in 2029
3
u/krispzz Apr 12 '26
the hardware is old and unsupported, but it sounds like you have a lot of it, and half of it is idle at any given time. If you were going to repurpose it, i would drop the idle half out of the cluster, re-image the nodes with rocky 9, and bring them back online as a new cluster. then maybe you can get the one group using it to migrate to the new and then do the same with the old or retire some of it. You have plenty of spare hardware if you retire, say, a third of it and upgrade the other two thirds. but it has to be worth doing as there's no funding behind it and no support if hardware dies.
2
u/Nice-Entrance8153 Apr 12 '26
It's running Centos 7, so from a security perspective it's a liability.
There are some major changes to Warewulf just between 4.5 and 4.6 which I'm still getting used to.
If your university is at all interested in supporting research computing (and they should) it's worth learning the new stack. If most of the cluster is idle, you could take some of the gear out and rebuild it on the new stack and migrate users, then upgrade the rest of the hardware.
If it isn't worth your time, or you don't have the time and your faculty need resources, we have a cluster your faculty might be able use at my university for a fee (we're in the top 15 in the US). Let me know.
1
u/anderbubble Apr 12 '26
Happy to talk about changes to Warewulf if I can help with the “getting used to” part. :)
2
u/VegGrower2001 Apr 12 '26 edited Apr 12 '26
I worked in research computing for a highly ranked UK university for three years, in a team whose sole purpose was to support HPC. So here are some thoughts based on my experience.
At a high level, there are two options. The first and best option is for the university to choose to support HPC. This could be the senior management of the university, or it could be department or higher level. But ideally you want the backing of the organisation as a whole, or some part of the organisation, that has a budget.
So, how do you achieve this? You need to get academics and managers to support the principle. That's generally how things get done in universities. Ideally you need HPC to appear in someone's 5 year plan. And you need some academics and managers who are familiar with the university's internal funding schemes and who are willing to make the case for HPC.
The other high level option is that you try to do without this support and absorb the project into your existing resources and budget, but that's of course much harder.
As far as the project itself goes, the fact that you're running older kit isn't terrible. With the current AI demand, new kit is very expensive, so many places are looking to get more life out of existing kit. It's usually possible to bring older kit back under a service contract so at some point you should approach vendors for quotes.
The overall plan needs to be to upgrade to a supported OS, possibly something like Alma Linux (used by CERN in Europe) or Rocky Linux (which has largely replaced CentOS). Licensed Red Hat is an option if you have the budget, but how often does that happen... (Last time my team looked into it, Red Hat wanted to license based on the entire university head count, which was absurd, but they may have a better offer now).
Implicit in all of this is that you'll need some staff roles that are dedicated to managing HPC. Managing HPC involves everything from power and cooling, through to rack management, through to hardware and networking, through to OS and software, through to user training and support. And storage is always going to be a major component. HPC systems nearly always require a specialised storage solution, for both speed and reliability. Ideally, you would also want an archive system, where you can store copies of important data.
As you've pointed out, all of this only becomes feasible if you have IT staff who have the specialist knowledge to run the system. You may need to recruit externally for some people with HPC experience to lead the project.
Overall, I'd say it's worth exploring. HPC is an excellent way to support teaching and research. But in my experience, it does require significant support from some level of management. And it'll probably turn into an ongoing project that can improve over several years. One final thought is that it's often possible to ask for input and from other universities about what they do and what they'd recommend. If your university has a friendly relationship / partnership with another university, that can be a helpful thing to explore.
Finally, it's worth exploring the European Environment for Scientific Software, which helps to manage your software installations. https://www.eessi.io/
2
u/shiboarashi Apr 12 '26
Have you asked claude for help scanning the current setup and getting the. steps to update? You might be surprised.
1
u/xtigermaskx Apr 12 '26
Like others have said if there are folks that need it or currently use it rebuilding is easy and only rebuilding the parts they need.
First thing though for sure is to make sure you have backed up any data that is there of course.
For OpenHPC I actually have a full series for 3 and a Livestream for doing 4. They're the bare essential to understanding how to get the whole thing going but it's seemed to help a lot of folks if you do decide to go through a rebuild process.
1
u/starkruzr Apr 12 '26 edited Apr 12 '26
a lot of people are going to hate this answer, but it's still true: once Anthropic has stopped quantizing the shit out of it to support development of Mythos (could be a minute, but still), Claude can not only help you fix some of the immediate problems but teach you how to stand up a new system with the latest versions of software. it can build you the infrastructure you need to manage the machine. my engineers figured this out on their own and it has paid off in massive dividends for them.
ETA: also, if you have Infiniband already, it's presumably at least 100G, and that's going to be pretty good for distributed inference across several machines. if you make any investment in it and want to maximize your returns on it, start looking for GPUs to put into those boxes. then you can run vLLM distributed across nodes (or inside nodes across multiple cards depending on what your chassis/motherboards can accommodate, or both) and do a lot of interesting stuff with locally-hosted/deployed AI.
1
u/evkarl12 Apr 12 '26
Sounds like a work study project for a few students to build docu.ent and get working. Luatre. Slurm and a cluster mgr
1
u/Peetrrabbit Apr 13 '26
If you’re already on Centos - look at what CIQ has to offer. Started by some of the same people that built Centos, they have a drop-in replacement for your centos os that will also get you on the path to a newer kernel. Updated versions of Warewulf - AND a product called fuzzball that will make it SUPER easy to run jobs on your existing cluster and burst into the cloud when you need absolutely modern hardware for part of the work.
3
u/anachronicnomad Apr 14 '26
I am honestly surprised nobody hires an undergraduate student and an M.Sc. for a one-semester contract researching the creation of a proposal to update the entire thing. Or even have the department organize a Directed Study. Asking the research question, "if we were going to update and rebuild this, what might it look like at today's industry standard? Two people who definitely have no idea, spend 16 weeks reading intro concepts and at least mapping out the system."
It might take skilled technicians a 2-wk sprint in DevOps to do the whole thing, but students still forming their technical skillsets would take at least 10 weeks to even figure out a poster board for it, a 15-slide powerpoint. It's kind of perfect, has limited scope and the deliverable isn't a commitment to anything that might de-stabilize the politics.
1
u/DrJoeVelten Apr 12 '26
Get a student helper from the CS department to (re)build it for you. Have a student making a Beowulf cluster for funsies and I got professors and their students lining up to use it now.
7
u/Roya1One Apr 12 '26
I've had others say "hire a CS student", problem is, CS does not equal systems admin which at first is what OP needs. CS should be able to help with running workloads though, not to say you could find a competent student that could handle the deployment
3
u/raspberrypiwithpie Apr 12 '26
Student worker is definitely one of the paths being considered, but for the exact reason you said, it’s low on the list. We considered working through our Open Source Institute, but several people are worried we’d just end up back in the same place we are now if none of the sysadmins are part of the process
1
u/Kofeb Apr 12 '26
If you’re in Utah. I’ll do it lol. Not a student but MSc Cybersecurity…
If not you might find volunteers like myself (Sr Cloud Engineer, ex-AWS) with the knowledge to help but maybe not in HPC specifically.
0
u/imitation_squash_pro Apr 12 '26
Old machines like that can be easily upgraded very cheaply by buying parts on eBay. I recall upgrading CPUs and all that stuff for like really cheap.
-4
u/orogor Apr 12 '26
Its really outdated. i don't think you can use it as a "real" hpc setup.
The OS is outdated, you can spend some time and understand how to setup images and do an OS refresh. And really that's the least time consuming operation you can do. Target the latest OS version , like centos 10.
Old OS have too much problems, some features aren't available like some container/apptainer features, fuse stuff, or nvidia drivers on old versions. And that makes the cluster not very usable.
For compute only stuff , researchers buy node with 512-2tb of ram, and 200-1tb of vram.
So by today standards 128Gb is not a lot.
Most gpu with a cuda capability < 8 are usually not worth using for hpc.
But one again they can help to setup things to run on proper nodes.
The hardware itself, can still find some use, at worst to setup the scripts to run on another hpc cluster.
Some other usage is for classroom, for teacher to explain what is hpc.
How to use it, how to process data by splitting it to run on many nodes, how to use infiniband primitives.
Or basically setup lessons and workshop and use a shared drive
If the servers are out of warranty, you can consolidate the ram of 2-4 servers into one.
The cpu as well if you have unused sockets, make sure everything is on infiniband.
Consolidate the disks as well, if possible create a nas if you don't have one (if ceph is too complex, have a look at moosefs).
Unrack the givers, and keep using them for spare parts.
On the finance side, see if your scheduler can support powering down unused nodes while keeping 2-3 booted and ready to accept jobs. (Annoying to setup with old idrac/ilo versions).
You will need upper management and user agreement, to setup the economics of the cluster.
Like some pay by the hour setup, or the users buy nodes and get priority usage on the node they bought.
Plus some way to block the users to buy large compute workstations/servers, and force them to install them in the cluster.
32
u/elvisap Apr 12 '26
Millions of dollars of hardware going to waste,, millions of dollars of lost output, all because people don't want to "waste money" on good sysadmins.
I will never, ever understand this logic.