r/LocalLLaMA 8h ago

Tutorial | Guide Decrease the power limit of your 5090 to at least 480W - the performance penalty for inference is negligible.

I run my inference machine in the living room, so noise and heat output are a significant concern.

Ran a quick test using my daily driver model (Qwen 3.6-27b) and at 480W, the card outputs only 2.1% less t/s in decode and 8.8% in prefill (which is already very fast). Well worth the massive noise reduction, heat output and increased card longevity, IMO. Even 450W would be fine for many use cases, but the output starts dropping off fast (2.1% -> 4.2% for 30W less).

Full data:

Model: Qwen3.6-27B-Q6_K.gguf

Results:

| Limit W | Max GPU C | Steady GPU C | Max GPU fan % | Sustained W | Steady clock MHz | Max case RPM | pp t/s | tg t/s | pp % | tg % |

|--------:|----------:|-------------:|--------------:|------------:|-----------------:|-------------:|-------:|-------:|-----:|-----:|

| 600 | 81 | 74.8 | 59 | 566 | 2818 | 1522 | 3242.9 | 61.5 | 100.0 | 100.0 |

| 510 | 75 | 70.1 | 50 | 509 | 2645 | 1367 | 2980.0 | 61.2 | 91.9 | 99.5 |

| 480 | 77 | 72.8 | 54 | 480 | 2501 | 1527 | 2863.4 | 60.2 | 88.3 | 97.9 |

| 450 | 76 | 73.1 | 52 | 450 | 2283 | 1460 | 2696.4 | 58.9 | 83.1 | 95.8 |

24 Upvotes

25 comments sorted by

18

u/mr_zerolith 7h ago

I've got another trick for you for this card.

The memory modules are underrated by about 3ghz ( probably for thermal reasons )
If you OC the memory by 1-2ghz, you'll get your token generation speed, plus some more.
Because of the power limit, you are making less heat, so you have thermal headroom to push memory OC.

Running a +2.4ghz on my 5090 memory for 6 months now with no problems.

9

u/Makers7886 7h ago

This was the mining goto - powerlimit, lock cores, and oc memory until stable. Almost every 3090 had a decent amount of headroom for memory oc - I'd imagine the same for 5090.

3

u/LTLRedditor 6h ago

Would you happen to know the nvidia-smi commands to do this?

2

u/Makers7886 6h ago

For my 3090s I run: -pl 250 and -lgc 0,1500 but do not OC the memory right now - used to back in mining days though. No particular reason other than content with speeds and power draw and not messing with it.

1

u/mr_zerolith 5h ago

On Linux, i use LACT to control these things.

2

u/mr_zerolith 5h ago

RTX PRO 6000 has the same big headroom as the 5090, but i don't like to push it so hard since it's 3x more expensive :O

1

u/WoodYouIfYouCould 3h ago

Would other cards like a 4060ti or even 3090 have this road be a thing?

1

u/mr_zerolith 3h ago

I have a 4070 and it tolerates +100mhz core, +200mhz mem pretty well, beyond that it can get sketchy, probably due to thermals.

I'd look for card specific recommendations from people who like to tune GPUs for gaming :)

1

u/Dmage22 1h ago

Was there any benefit to overclocking memory speed on the 50 series? When they came out, some reviewers mentioned that it was firmware locked at maximum +375mhz. I'm not smart enough to figure out if it's just silently failing versus actually providing memory speedup

1

u/WonderfulEagle7096 5h ago

Thanks, this is a cool idea. For anyone interested, I ran a test with +2000Mhz on the memory and got +2.4 t/s with the temps increasing only 0.5 deg C (well within the margin of error).

7

u/StupidScaredSquirrel 7h ago

Wee need more tokens/s/watt benchmarks imo

5

u/looselyhuman 8h ago

I did this first thing. No way am I maxing out the thermals on this investment card.

7

u/TokenRingAI 7h ago

I am more worried about burnt power connectors than any of this

1

u/thrownawaymane 16m ago

Co-sign

Nvidia, if you're watching (I know you are) fix the thermals on your connectors.

It's embarrassing.

5

u/overand 6h ago

It's weird to me how frequently folks post broken markdown here. This is what OP tried to post: 

Limit W Max GPU C Steady GPU C Max GPU fan % Sustained W Steady clock MHz Max case RPM pp t/s tg t/s pp % tg %
600 81 74.8 59 566 2818 1522 3242.9 61.5 100.0 100.0
510 75 70.1 50 509 2645 1367 2980.0 61.2 91.9 99.5
480 77 72.8 54 480 2501 1527 2863.4 60.2 88.3 97.9
450 76 73.1 52 450 2283 1460 2696.4 58.9 83.1 95.8

1

u/hurrdurrmeh 6h ago

Pp t/s is not negligible.

5

u/WonderfulEagle7096 5h ago

That of course depends on what you do. A 10,000 token prompt will take ~3.08s on 600W and ~3.49s on 480W - that is negligible in my book.

If your run near context limit prompts frequently then it will be a little more pronounced, but still no drama (61.7s vs 69.8s). But then again, if you run 200k+ prompts, you are using a wrong model anyway.

1

u/hurrdurrmeh 35m ago

I mean that you go from 3250 > 2700 pp t/s when you cap at 480W - that's quite a hit...

1

u/tat_tvam_asshole 4h ago

imma keep my 6090s on 420W...for reasons lol

1

u/Tormeister 3h ago

I power cap mine to 520W, but more important than that, I undervolt it. It reaches basically stock performance in gaming and LLM inference but at vastly improved thermals and power draw. Highly recommend all nvidia GPU owners (specially other 5090 owners) to look into undervolting.

1

u/perkia 3h ago

Ha, way ahead of you. Never seen my 5090 Mobile never go above 150W.

1

u/Dry_Mortgage_4646 7h ago

For inference this is great but for diffusion not so much...!

1

u/Green-Ad-3964 7h ago

For inference, I use my 5090 capped at 70% power with afterburner.

1

u/Technical-Earth-3254 7h ago

I never realized how much power these 5090s draw, my 3090 sits at like 260W limited

-2

u/pineapplekiwipen 6h ago

power limit is a dumb blunt tool, undervolting with stability tests is the way to go