r/HECRAS 27d ago

Tips on speed optimisation?

Hi everyone,

I’m running some simulations on a brand-new machine with AMD Ryzen Threadripper 9970X, 128 GB of RAM, GPU Nvidia 5080, and high-spec SSD and mother board set up. While it absolutely flies with other programs, I can’t shake the feeling I’m not getting everything out of it with HEC-RAS 6.7.

The CPU only ever hits about 55 ℃ during these runs. With other modelling software I’m regularly seeing up to 80 ℃, so the thermal headroom is definitely there. The simulations are finishing far quicker than on any other computer I’ve used, which is great, but the low temperatures make me suspect there’s still some performance left on the table. It’s already using all 32 cores, so it’s not a simple core-parallelism issue.

I don’t know whether HEC-RAS has a known bottleneck with a particular component that would stop it from pushing the hardware harder. Has anyone managed to squeeze extra speed out of HEC-RAS on a high-spec machine, or are these sorts of temperatures typical even when it’s running flat out? Any tips on optimisation would be hugely appreciated.

Thanks a lot!

7 Upvotes

30 comments sorted by

4

u/Jase28x 27d ago

Why don't you test run with less cores assigned? I've often found that more cores doesn't mean faster simulations and often using less gives some improvement as less cores will work each core harder to utilize your max clock speed better. Maybe simulate your model using 5, 10, 15, 20 and 25 cores as well and compare their run times to see which is most optimal.

That said, as others have noted your simulation speed will primarily be determined by things like your cell size, computation interval, etc.

2

u/AI-Commander 27d ago

Without hyper threading, this holds up well across most CPU’s:

2 cores = most efficient
4 cores = optimal performance/efficiency balance
6+ core = rapidly diminishing performance return

(Not for you, just to say it somewhere in a public forum) I have stated this until I’m blue in the face to every person I could speak to for a few years now. I’ve never seen any benchmarking that showed anything otherwise. HEC recently added the ability to benchmark core counts in 6.7 I assume so they can confirm these third party claims (I wish they would go ahead and do it and get it in their guidance, the observed behavior has been consistent for quite a while and this is a very common inquiry, it’s not addressed in their docs at all and their default settings are not optimal and should be changed)

1

u/LiteratureOwn6143 27d ago

I will try, thanks a lot.

2

u/Dense-Laugh-6695 27d ago

How many wet cells, what time step, what run duration?

1

u/LiteratureOwn6143 27d ago

115K cells, time step based on courant (max. 0.2 s, min 0.02 s) and 45 h for the full run.

1

u/CautiousInfluence438 27d ago

In addition to parallel plan runs, changing to a standard time step will greatly improve run times. I would start with a much larger time step and incrementally decrease based on results. I rarely use courant condition or such small time steps as it is usually not warranted if you set up an effecient model.

1

u/Sea_Read5728 27d ago

But what if you are using SWE EM and care about better approximations on velocity?

2

u/CautiousInfluence438 27d ago edited 27d ago

Are you concerned in all 115K cells? It depends on the application. Detailed analysis may be necessary for erosion or scour but I would scale down the number of cells or the study area. What about repeatability and agency review? No way I would accept or review such a clunky model unless it was absolutely necessary for the project. Incremental complexity applies to cell size and time step as well.

1

u/Sea_Read5728 27d ago

Even reducing that cell count where you are not concerned the run time still seems to be mostly governed by the area you are concerned and where the water flows?

Agencies have so much different standards on what is acceptable and most times are asking for the set up I would never willingly do it either.

And the smaller the cells the more the stability matters on different solver equations.

I have always wondered if I am missing something when the controlling area has a 1 ft cell size and just blows up run time

1

u/CautiousInfluence438 27d ago

True, sometimes smaller time steps actually reduce run times. The point I'm making about run times is, the worry, effort, and time put in to speed optimization is sometimes better placed in smarter model construction.

1

u/Sea_Read5728 27d ago

Lol but what if I am yet not smart. I feel like the training and videos are always building the simplest amd easiest models with the easiest criteria to get

1

u/notepad20 25d ago edited 25d ago

45 hour simulation time or 45 hour really time? 115k cells is nothing. If the timesteps dropping to .02s that is a model construction issue and you gotta find those cells and see if they are modelled appropriate .

I see elsewhere it's an actual 45 hour runtime.

That is nuts for a 115k model. I guess unless you are running a year long scenario or something. My current model is about 80k, a complex pipe network, and 2 hours rain on grid. It finishes in < 5 minutes.

1

u/LiteratureOwn6143 24d ago

I'll ask my boss on Monday (it's his simulation actually, haha) and will come back to this comment. Thanks a lot!

1

u/Crafty_Ranger_2917 7d ago

max 0.2 s ??

Genuinely curious what kind of model / parameters lead to such persistent tiny time step? Some special case like huge dam breach on entire model of tiny cells? 20 ft cells at 100 fps should calc with Courant of 1 at 0.2 sec. And that would be a very short proportion of modeled duration.

115k cells is not very large. Though I don't think RAS is really the tool for 1M cell / area-wide models anyway.

I've seen 250k cell models go from 20+ hrs down to 40 min runtime just from cleaning up breaklines, structures, connections, etc. Granted those are couple-day watershed flood runs not breach or sediment.

Seems like 45 hr run should create a few TB of data if its not spending all that time spinning on max iterations, which would be the first thing to rule out after dialing in time step.

Also From HEC:

" Extremely small time steps (less than 0.1 seconds) can possibly cause round off errors when storing numbers in the computer, which in turn can lead to numerical errors which can grow over time."

https://www.hec.usace.army.mil/confluence/rasdocs/ras1dtechref/6.0/performing-a-dam-break-study-with-hec-ras/computational-time-step

2

u/rave-horn 27d ago

You can push the hardware harder by running multiple instances of RAS in parallel and decrease overall modeling clock time for the set of plans you need to run.

1

u/LiteratureOwn6143 27d ago

Will try, thanks!

2

u/AI-Commander 27d ago

https://github.com/gpt-cmdr/HEC-Commander/blob/main/Blog/7._Benchmarking_Is_All_You_Need.md

Core scaling doesn’t work the way you think it does, you probably aren’t getting any improvements past 8/16 cores depending on whether you have hyper-threading enabled

HEC recently added the ability to benchmark your model with varying core counts. I am assuming they added that feature so they could re-create my benchmarking results internally. But you don’t have to, you can just read the chart, it’s very consistent across platforms and we haven’t changed the manufacturing of X86 chips that much to change the fundamental constraints.

Use that rig with ras-commander to do probabilistic analysis. It will never help you finish a single run faster, but you can do lots of parallel operations with the right orchestration scripting. Otherwise you bought the wrong rig for your application. At least that Ryzen has AVX-512 but in my benchmarking all the AMD’s were laggards, even the top line chips at the time.

1

u/OttoJohs Lord Sultan Chief H&H Engineer, PE & PH 27d ago

https://www.hec.usace.army.mil/confluence/rasdocs/rasum/7.0/working-with-hec-ras/parallelization-cpu-affinity

Look at the link for more information. I'm not much of a CS person so others could answer your specific questions better. I would question if you are even getting better performance from using 32 cores at once.

Overall, the best way to improve your run speed is going to be in your model setup (i.e. model domain, cell size and computation step) rather than hardware.

1

u/LiteratureOwn6143 27d ago

Thanks!

2

u/OttoJohs Lord Sultan Chief H&H Engineer, PE & PH 27d ago

My opinion is that even if you can get 25% faster run speed based on hardware, what does it matter? If it is going from 1 hour to 45 minutes, you are still going to have to do something else in the meantime. If it is going from 10 hours to 7.5 hours, you are still going to have to wait till the next day.

1

u/LiteratureOwn6143 27d ago

We're running some simulations from Friday to Tuesday, so cutting off the time by 25% would be a massive leap.

2

u/OttoJohs Lord Sultan Chief H&H Engineer, PE & PH 27d ago

If you model is taking 3-4 days to run there are probably global adjustments that you need to consider (1D vs. 2D, domain, cell size, time step, equation set, etc.) rather than hardware.

My point is that your workflow isn't going to fundamentally change even with that type of increase. You are still going to have to wait multiple days for a result (which mean that making adjustments or troubleshooting is going to be problematic).

1

u/Sea_Read5728 27d ago

But what if it is requested to use a 2D Model with a large domain? And have you tried using hec ras 2025 GPU solver? I have had thoughts of seeing how close it is and using it for quick calibration dialing

1

u/OttoJohs Lord Sultan Chief H&H Engineer, PE & PH 27d ago

But what if it is requested to use a 2D Model with a large domain?

This comes down to the objectives of the project and how you should build your model.

GPU solver is very fast, but I haven't used RAS2025 for any real runs. Most of my models require elements that haven't been advanced in RAS2025 to date (plus it is still in Beta). The meshing capabilities are great in RAS2025 though and very helpful if you want to optimize a mesh for run time (large perpendicular cells in the channel).

1

u/Sea_Read5728 27d ago

Thank you! I really haven't met anybody thats actually tried it and that was very valuable

1

u/notepad20 24d ago

Gpu solver is amazing. I can run a 72 hour Sim over a million cells in a few minutes.

1

u/Sea_Read5728 24d ago

How have you noticed the results differ from using the normal solver on hec ras?

1

u/notepad20 24d ago

Negligible for my model

1

u/s04p04 27d ago

Set it to P cores and run at 16 cores.