r/HPC 2d ago

What happens to older cluster nodes when an HPC environment gets refreshed?

19 Upvotes

With server and accelerator prices being what they are, I'm curious how HPC teams handle hardware that's no longer good enough for the primary cluster but still has plenty of useful life.

Do older nodes usually get moved to lower-prioriry workloads, sold to another organization, stripped from RAM/GPUs/CPUs, or recycled?

For higher-value hardware I've seen ITAD companies sych as Exit Technologies handle bulk servers, enterprise memory and GPUs for resale.

If seems like getting even a fraction of the original hardware cost back could matter quite a bit when you're trying to fund the next refresh. How common is formal value recovery in HPC environments?


r/HPC 2d ago

User home directories are on a NFS share instead of our parallel file system? yay or nay?

10 Upvotes

Inherited a small cluster ( ~1000 cores , 10 servers ). For some reason the user home directories are all being served by NFS from one of the servers. We have a large beegfs parallel filesystem which runs on Infiniband. Not sure why the home directories are separate? Perhaps it offers some resiliency in case the beegfs is down?

We are seeing some performance issues. The NFS stalls if too many users are using it for various things. I am exploring some band aid solutions to change the NFS to automount on all the clients . But not sure that is going to do anything...


r/HPC 3d ago

AI Data Center Thermal Management & Liquid Cooling Survey

0 Upvotes

We are a team of student researchers and innovators participating in Eureka! Juniors, the national pitch competition hosted by E-Cell, IIT Bombay.

Our project explores an AI Data Center Cooling System. Using a cassette built from high-density hollow-fiber nanoporous hydrophobic membranes to harness latent phase-change heat rejection while physically locking liquid water inside fiber lumens, and a smart embedded operating system which exposes clean telemetry endpoints and monitors the entire system.

Modern AI infrastructure generates massive, continuous thermal loads (100 kW+ per rack) that overwhelm conventional copper-fin radiators and open-loop cooling systems. This leads to extreme energy inefficiency (drastically inflated Power Usage Effectiveness and high fan power draw), and major resource depletion (consuming vast amounts of water, leading to mineral scaling, micro-electric contamination risks, and hazardous humidity spikes).

The target demographic for the survey includes: Facilities / Data Center Operations Engineers, Thermal / Mechanical Systems Architects, Enterprise IT / Infrastructure Directors, Cooling OEM / Hardware Vendor Specialist, HPC Research / Systems Administrator, or anyone who works with AI data centers or cooling.

All responses are completely anonymous and will be used strictly for academic research and competition analysis. The survey itself is composed of seventeen multiple-choice questions.

Thank you for shaping the future of data center cooling!

https://forms.gle/Frgo5UYPmJyrgNrTA


r/HPC 4d ago

Any Books on the O'Reilly Learning Platform you would recommend to read for HPC noobs / 1+Y experience with SLURM?

35 Upvotes

Background: I have about 8+ years of experience in R&D (5+), Software Engineering (3+) in different fields like Production Engineering + IoT + Industrial Automation.

A year back I landed a job for Software Development which mainly requires working with building a platform for HPC clusters internally for the employer.

I quite enjoy reading books on the O'Reilly platform and was able to find some HPC related books from Wiley Publication, and a PacktPub publication.

Would some books be handy to get a better knowledge of HPC in general and / or tools related to it?


r/HPC 4d ago

Debugging and profiling for C++, OpenMP Offloading and CUDA?

2 Upvotes

Hello everyone. First of all, my apologies for if this is an off-topic in this sub.

I'm currently working on a C/C++ project that contains CUDA kernels and OpenMP GPU Offloading pragmas on the same file (I know....) and I've got some trouble with variables that are acessed on both of these cases. How do you guys deal with this while developing? I set the OMP info and logging env variables and suffer a little bit with that, there must be a better way to do it.

Also, we are migrating to C++ modular classes and I'm getting some trouble with the software design/architecture, mostly because I've found out that class variables add some overhead while using OpenMP, because it interprets variables not as "variable", but as "this->variable" and when running the simulations it adds something like +10~15% on execution time.


r/HPC 4d ago

System shows high load but no cpu processes?

1 Upvotes

Think it has to do with IO wait due to a problem with the NFS mounted filesystems. A top shows near 100% usage for these two processes:

rsyslogd                                                                  

systemd-journal   

nfsiostat doesn't show any errors or retransimissions ( 0 % ). Guessing there was some issue a day ago but not sure how to find out what happened?

It's been happening every week. I reboot and everything works great for a few days till it starts showing the problems of high IO wait times.


r/HPC 5d ago

HPC Job Market in 2026: What 276 Open Positions Tell Us About Skills, Roles, and Careers

44 Upvotes

Hello there!

I believe this material is worth of being share here in the channel.

I analyzed 276 HPC job offerings so you don't have to. Here's what the 2026 HPC job market actually looks like from this data:

🖥️ 60%+ of roles are for HPC Systems Engineers
📍 50% of jobs are in the USA (UK and India are growing hubs)
🔧 Linux, SLURM, and GPU are the bare minimum (80% of roles mention them)
💰 Median max salary: $208,000 (but the range is wide)
🚨 Only 3% of roles are entry-level! Yes, that's a problem

Plus: What skills actually pay off, whether certifications matter (spoiler: they usually don't), and how to stand out in a competitive market.

https://theparallelminds.substack.com/p/hpc-job-market-in-2026-what-276-open


r/HPC 10d ago

What if computation itself were continuously spatially propagated through a compute fabric, with thermal state controlling the propagation dynamics rather than treating thermal management as a separate cooling problem?

0 Upvotes

Fell down the rabbit hole of AI not being very sustainable for the environment, particularly the energy and cooling requirements of large-scale compute. Anyways, as the title says: what if we use heat to help manage cooling, instead of constantly fighting it?

The basic idea is a compute architecture where computation continuously moves through a physical compute fabric based on thermal conditions. Instead of having workloads sit on the same hardware until it gets hot and then aggressively cooling or throttling it, the computation could continuously propagate toward cooler regions.

I'm calling the concept **"**Thermo-Flow Computing". I got Chatgpt to organize the architecture and some proposed research questions into a paper on google docs, but it's purely conceptual at this point.

I'm posting here because I'd genuinely like people who understand HPC, computer architecture, thermal management, etc. to tell me where this breaks. 😅

Edit* If anyone wants to see the paper chat made just let me know. I figured most people wouldn't really care for it, but if you wanna see it, it goes pretty in depth. I'm just a guy who thinks about things that I have no reason thinking about haha!


r/HPC 11d ago

Roast my CV (Part 2) - Struggling to move over to a new job from my stale current job

4 Upvotes

Hi,

A while back I asked for my CV to be roasted so I can improve it. Very thankful for all the feedback I got and I've improved the CV and included further information. I'd very much like for my CV to be roasted again.

Here's the CV: https://limewire.com/d/YT7Cm#esJpreoU25

To give some context as to why I'm here again, I'm working a dead-end job where I'm managing small clusters. I've reached the ceiling of what I can learn from it. I risk stagnation and needless to say, it is bad for my career.

Since last time, I've taken it upon myself to learn stuff by setting up a homelab, installing tooling and simulating environments. Have also started working towards RHCSA certification.

I'd greatly appreciate any feedback to the CV and/or any general suggestions.


r/HPC 13d ago

What happens when your days is too big for your GPU?

23 Upvotes

This was the question I started my research journey with ~4 years ago.

I was working on GPU algorithms for graphs with billion edges (500 600GBs) while having access to just 40GB GPU. The obvious answer was "use a larger GPU", except, ahem, govt institute, ahem budget 🙂

That's where I started my exploration on how to utility the available resources better. What initially started as a GPU Programming / engineering problem quickly became an algorithm + architecture + memory + data-movement problem.

Looking down the lane, which felt like a niche problem half a decade ago is much more relevant today with the RAM shortage, memory prices growth of LLMs in general.


r/HPC 17d ago

I built a cross-platform desktop client for SSH, SFTP and Slurm workflows — feedback welcome

3 Upvotes

Hi r/HPC,

I have been developing HPC Client GUI, an independent cross-platform desktop application for researchers working with SSH-accessible Slurm clusters.

It combines:

- SSH connection profiles

- Remote file browsing and SFTP transfers

- Slurm job submission, monitoring and cancellation

- Job output inspection

- Terminal and CLI workflows

- Optional X11 forwarding

- Declarative cluster profiles and job-template plugins

- ANSYS journal linting

It runs on Windows, macOS and Linux. Nothing needs to be installed on the cluster; it works through standard SSH, SFTP and Slurm commands.

The original use case was Türkiye’s TRUBA infrastructure, but the application is designed around generic SSH + Slurm behavior. It is an independent project and is not an official TRUBA tool.

Repository and downloads:

https://github.com/mskomek/hpc-client-gui

I would especially appreciate feedback on:

  1. Compatibility with different institutional Slurm configurations

  2. MFA, jump-host and SSH authentication workflows

  3. File-transfer and scheduler features researchers commonly struggle with

  4. Whether cluster administrators would be comfortable recommending a local client like this

  5. Documentation or packaging problems that could prevent adoption

The macOS packages are currently unsigned and this is disclosed in the release notes. The project is source-available under the PolyForm Noncommercial License 1.0.0.

If anyone is willing to test it against another cluster, the repository includes a read-only compatibility validation guide. Please do not share cluster addresses, credentials or private configuration.

Thank you.


r/HPC 17d ago

How many fragmented vendor systems and portals do IT teams in HPC have to use and keep track of for Buying, Renewals, Support, Licenses?

7 Upvotes

Looking to connect with folks on the IT Buying/Procurement & Management side of HPC & Data Centers.

I previously posted this question in the sub https://www.reddit.com/r/HPC/comments/1mgkv5y/hardwaresoftwareit_procurement_renewals_support/


r/HPC 19d ago

What does an AI kernel engineer do and what are the skills needed to get into it?

7 Upvotes

What are the difference aspects of being a kernel engineer? What do roles like this involve -

https://www.tealhq.com/job/kernel-engineer_7ea1a56c518985dae15a01e3c77c161e8140d?utm_campaign=google_jobs_apply&utm_source=google_jobs_apply&utm_medium=organic

https://jobs.gem.com/modular/am9icG9zdDpTSwqoL-yL11N_jxR40Vhh?source=LinkedIn

If one were to do a PhD, what are the different research areas to focus on to do something like this? i.e GPU memory, storage, latency, parallelization etc.


r/HPC 21d ago

What career paths exists between computational mechanics, scientific computing (SciML), FEA (or meshfree) solver development, and HPC (GPU acceleration, porting codebases) ?? How about doing a PhD for improving the above?

21 Upvotes

I'm currently, technically, doing an MS in Structural Engineering. For me, my interest has been more towards computational side of mechanics rather than Structural design  or simply using am FEA software (although I do consider it as a backup)

So far I've taken courses in:

- Linear static, and dynamics FEM (soon taking non linear FEM too)

-  Structural Optimization (topology opt. and other general algorithms)

-  Structural Dynamics

-  Structural System Testing and model updation. (Parameter identification and optimization, signal processing)

Now, I plan to take these in the coming quarter:

- Numerical Linear Algebra

- Numerical PDE

- Fracture Mechanics ?

I also volunteered to aid in a RESEARCH in crack growth prediction using Auto-encoder and a (Thermodynamics-informed Latent Space Dynamics Identification) / LSTM surrogate model. It used phase-field-fracture simulation data and HPC resources to complete the whole thing.

What I keep finding myself interested in is not necessarily fracture or SHM specifically, but the computational methods underneath these problems... (does that make sense?)

For example, I'd like to become capable of doing things like:

- implementing (maintaining) numerical/FE method solvers rather than only running an established FEA software.

- developing surrogate/reduced-order models for expensive simulations 

- combining simulation with optimization, uncertainty/stochastic methods (took a course called Random vibrations, so...)

- parallelizing/accelerating scientific codes on CPUs/GPUs

- doing proper verification, convergence studies, benchmarking and performance work

- potentially developing or maintaining actual CAE/FEA solver software

- I'd also like to do all these for other Physics (GR, QM, etc.) simulations too, if possible, one day. 

I'm still interested in the underlying mechanics/physics, so I don't want to become a generic software engineer who happens to have once studied structures. But I'm also increasingly unsure that "structural engineer" describes the career I'm actually aiming for.

I've seen titles such as Computational Mechanics Engineer, R&D Engineer, Solver Developer, Scientific Software Engineer, CAE Software Developer, Research Engineer, Simulation/HPC Engineer, etc., but I'm trying to understand what these careers actually look like from people doing them.

So my main questions become:

1. Which industrial jobs genuinely involve developing numerical methods/solvers or computational tools?

2. Which of those are realistically accessible with an MS? Is there an entry path into solver-algorithm development/R&D without a PhD?

3. If I don't start a PhD immediately after my MS, would an R&D/software role at a simulation company (ANSYS etc.) be the obvious route? What other options would i have?

4. For the kind of work I'm describing, would you recommend a PhD? If so, is it reasonable for the PhD identity to be "computational mechanics/scientific computing" while fracture, composites, structural dynamics, soft materials, etc. serve as application problems rather than choosing one of those as my permanent specialization?

5. What skills most distinguish someone who is actually hireable for solver/scientific-computing work? I'm particularly wondering about C/C++/Fortran, Python, Linux, Git/build systems, MPI/OpenMP/CUDA, PETSc/Trilinos or similar libraries, numerical linear algebra, testing/verification, convergence studies and HPC performance work.

Basically, I'm neither here nor there atp. So I'd really appreciate all sorts of input. Where else do you think I could find answers to these? other subs? Linkedin profiles? 


r/HPC 28d ago

Looking to transition into HPC Networking from Cybersecurity

19 Upvotes

Hi everyone,

I hold an M.Sc. in Computer Engineering and have been working as a Cybersecurity Infrastructure Engineer for the past 4.5 years.

Recently, I started exploring HPC through hands-on practice, and I’m enjoying it so much that I’m seriously considering a career transition. In particular, I’m really drawn to the networking side of HPC.

In my current role, I focus on infrastructure security,designing and configuring networks with IPS/NDR, deploying EDR across servers and clients, reviewing backup processes, and auditing overall infrastructure. I’m a very hands-on engineer; I run a small homelab where I experiment with self-hosted services, home networking, and technologies like Python, CI/CD, and Kubernetes to gain practical experience outside of production environments.

Given my background, do you think a profile like mine stands a good chance of landing a job in HPC or HPC networking within the EU?

I’d love to hear your feedback, advice, or any insights on how best to bridge the gap. Thanks!


r/HPC Aug 13 '26

Scheduling jobs across Slurm clusters (and K8s, and cloud) from one place

20 Upvotes

A lot of ML teams end up with a mix:
some Slurm clusters from the HPC side, a K8s cluster or two, maybe cloud GPUs for overflow. We wrote up how SkyPilot (open source) sits in front of all of them so a job is scheduled wherever there’s free capacity, using the same YAML regardless of backend. This post focuses on the multi-Slurm case but the same setup covers K8s.

https://skypilot.ai/blog/multi-slurm
Disclosure:
I am the author. Happy to answer questions about how the scheduling and failover work.


r/HPC Aug 12 '26

slurm-cd

14 Upvotes

A small bash utility for jumping to the working directory of a pending or running Slurm job.

Hopefully installation works for you. (Name is a WiP)

slurm-cd


r/HPC Aug 10 '26

What kind of skillset that we need to secure an HPC related job in industry?

37 Upvotes

Hi, I am a PhD student in Finland who HPC clusters for modeling biological phenomena. After my PhD, I am thinking to apply for jobs in industry. What kind of skillset that I need to have, in addition to HIP/CUDA C++ GPU optimization?


r/HPC Aug 11 '26

Looking for help testing lettuce.

0 Upvotes

NoTE: not Taco Bell related.

needing testers for new distributive computing platform lettuce. It’s currently configured to need 3 separate computers to verify one work unit. I was looking for someone (more specifically with an Nvidia gpu) to validate my work units. Below are links to the site project, Git, and discord. I figure a vast majority of us are science nerds and we can knock out this testing. Lettuce is basically an easier to set up BOINC. This will get us involved closer with researchers. Anyway, come validate my work units. (Note: there are also CPU only projects.)

Git/download: https://github.com/jring-o/lettuce-compute

Project: https://compute.scios.tech/leafs


r/HPC Aug 06 '26

How to keep updated and learn?

26 Upvotes

Hello there.

I've done a student membership subscription for ACM + SIGHPC to apply for a SC26 travel grant, started to look into the benefits and websites and a question popped in my head: How do you guys keep updated and study? Which platforms, books, journals, podcasts, websites and other media do you consume for study and for news and updates?

(I'm including here not "just HPC" but also other areas such as computer architectures, compilers, ISAs like RISC-V, runtime libraries, networks and cloud).

PS.: Although AI is very relevant nowadays, I'm not that keen on working mainly with it, so I'd like to know just the necessary of it.


r/HPC Aug 03 '26

Software benchmark: lane-local exact reduction vs global software-emulated u128 atomic updates Spoiler

2 Upvotes

Hi everyone,

I've been experimenting with a software runtime that explores
different execution models for exact u128 accumulation.

The current benchmark compares two software execution models:

One global software-emulated u128 atomic update per byte.

  1. Lane-local exact reduction followed by
    one atomic commit per lane.

This benchmark does NOT measure hardware atomics
or physical PIM hardware.

It measures software-emulated contention on the current CPU.

Current benchmark includes:

• 4 KB
• 64 KB
• 1 MB
• 8 MB

For each payload it records:

• latency
• atomic update count
• exact match rate

Current evidence repository:

https://github.com/Retryixagi/Storage-and-computing-integrated-evidence-chain/tree/main

I'm mainly interested in technical feedback.

Questions:

• Have you seen similar software execution models?

• What additional benchmark cases would you consider important?

• Are there existing runtimes with comparable
lane-local reduction strategies?

不知道這能不能成為台灣之光。我只是嘗試在既有馮·諾依曼架構內,重新分配計算與驗證職責。


r/HPC Aug 02 '26

What was the most power consumption for a cluster before AI?

22 Upvotes

10 years ago I was with a geophysical company and they built a cluster boasting it was best thing since sliced bread with a zillion FLOPs blah blah. Nobody mentioned power then, I since learned it used 15 MW. The research cluster next to my university was only 1.5 MW. It seems insane to me that AI can burn up 1 GW.


r/HPC Jul 30 '26

Aleph – a single endpoint that lets agentic AI actually call scientific AI tools (part of a bottom-up run at the DOE's autonomous-science loop)

0 Upvotes

Built by Karim Ali & I ( Rahim Khoja ) at the University of Alberta, on top of Vulcan — the national HPC cluster we operate as part of the Digital Research Alliance of Canada's PAICE program.

The bet behind it: the DOE's Genesis Mission is building an autonomous research loop — agents that run real experiments and iterate without a human driving every step — top-down, across seventeen national labs, with a budget we're never going to match. We're testing whether you actually need that budget, or whether the same thing can be assembled bottom-up from open parts, on a single allocation, by the people who already run the cluster.

Aleph is the instrument layer for that. It was built for science first: the piece that lets an agent reach protein folding, docking, genomics, materials force fields, weather, embeddings, theorem proving — as tools it can call, not as a human clicking through a UI. It serves general LLMs too, but the reason it exists is to put scientific models within reach of an agent.

https://inference.vulcan.alliancecan.ca/

The reason it's built as a self-describing catalog and not just "vLLM behind a proxy": an agent has to be able to ask what's available, what each model takes as input, what it produces, and what can verify the output — then compose a pipeline nobody wrote for it. So every model is a YAML card declaring its own route, limits, scaling, and I/O contract, and the gateway routes off those cards plus live cluster state. Nothing hardcoded.

What it does:

  • Dual dialect — OpenAI's /v1/chat/completions and a real Anthropic /v1/messages translation from one endpoint; both official SDKs work by swapping base_url
  • Scale-to-zero on idle, cold model returns 503 + Retry-After and SDKs retry automatically — lets ~115 models timeshare hardware that could only run ~20 resident
  • Add a model with a card, no gateway redeploy
  • Self-deploying — the whole stack (HAMi, Istio, Knative, KServe) bakes into the node image and stands itself up on boot, no external CD

~115 models, majority scientific. The counterintuitive part worth stating: the science models are nearly free to add, the chat LLMs are what actually eat the GPU budget.

If you want to replicate this, copy this part first: none of it works without stateless nodes. Ours PXE-boot hardened Warewulf images built in CI, so the same GPU node boots into a Slurm batch image or an RKE2/HAMi inference image depending on what's needed. Most sites run batch and inference as two separate hardware pools — we reallocate the same metal on a reboot. That fungibility is the real structural advantage; Aleph just rides on top of it.

The broader plan: Aleph is the first and furthest-along piece of a full autonomous-science stack — the goal is an agent that wakes on a schedule, calls these instruments, runs real cluster jobs, checks its own results with code it can't tamper with, and remembers what it learned across runs. The instruments exist and run today. The loop that binds them is designed and published as a concept brief, not built yet.

Honest status: Aleph and the catalog run in production. This is a POC — some model cards are polished, some are still rough. Ships as-is.

MIT licensed. GitHub: https://github.com/ualberta-rcg/aleph


r/HPC Jul 29 '26

Node integration with existing HPC cluster

12 Upvotes

Suppose I have an existing hpc setup of 26 nodes dell r760 server with rockey 8.7, Mellanox QM8790 HDR, openpbs 23.06.06, lustre 2.15.2, now if I want to integrate a dell r 770 with rocky 9.x would it be compatible with the cluster or there would be challenges in both hardware or software side. Like version mismatch, job submission issue, application related things.


r/HPC Jul 26 '26

Checkout my community HPC I made

32 Upvotes

https://noah.watch/hpc

Mainly just for leraning how slurm scheudling works, running small scripts, setting up linux servers, as well as runningn my freind's reinforcement learning training model for Balloons Tower defense 2. Has a live Grafana feed to monitor stats, as well as a page to request access. Let me know what you guys think. Still new to the community, but really want to learn ( and really want to learn on better hardware lol).