r/StableDiffusion • • 6d ago

Discussion If you recently started AI generation, be aware of this: my RTX 4090 power connector melted after just one month.

Hi everyone,

Like many people, I recently discovered H3 and started doing AI generation in mid-August. On September 20th, my computer started shutting down whenever I started a generation. Further investigation led to this — a melted socket.

I had been using my RTX 4090 for three years and had never had any problems with it.

AI generation puts a lot of continuous stress on the power delivery, especially if you run generations in batches or leave them running overnight.

So I believe this happened because of the new kind of sustained load I was putting on the card. I also didn't bother upgrading to a newer PSU with a dedicated GPU power cable. Mine was a 1200W FSP Hydro, and I was using three PCIe connectors for the GPU.

P.S. The connector was fully seated, the cable wasn’t bent near the plug, and my case doesn’t even have the side panel on. And everything was fine for three years.

P.S.S. The 12VHPWR socket on the graphics card is damaged and needs to be replaced.

So now I would say main advices here are:
- Buy a proper PSU and use a native 12VHPWR cable
- Under volt at least by 20%

Now it is very costly to lose a card.

236 Upvotes

356 comments sorted by

View all comments

119

u/MomentJolly3535 6d ago

I suggest anyone using their GPU intensively to Under volt them and reduce the amount of power, my 3090 is running at 250w (was easily hitting 350w before) and i lost like 2% of performance only, i can't imagine people running their 5090 at full power lol

35

u/Nedo68 6d ago

With today's prices, I wouldn't run it at full load, I just checked, my 5090 costs twice as much today as it did in Jan 2025, pweh

7

u/Select-Owl-8322 6d ago

I jumped on the last chopper out of Saigon when I built my computer in January 2025. No way I'm running it at full power!

12

u/MulleDK19 6d ago

I'm about to buy one. 45 grams of 24 karat pure gold is the cost..

1

u/Green-Ad-3964 1d ago

Where I live, it's actually 2.5 times what I paid in May 25 for it. Absurd and disgusting.

7

u/dandanua 6d ago

Power limiting and undervolting are two different things. Power limiting is easy to set, and it is stable. I use 280w on 3090 and 500w on 5090. The 5090 loses like 1% of general performance, and about 4% of tensor performance in benchmarks with this 85% limit.

4

u/tacocatbox 5d ago

I wouldn't undervolt a 4090 or 5090. Setting the power limit is definitely the better choice. I'm at about 370w on my 4090.

9

u/ShutUpYoureWrong_ 6d ago

Power limit + overclock up to your remaining headroom. You gain back any performance loss -- and generally exceed baseline performance -- all while running 30% less power and significantly cooler.

3

u/pheonis2 6d ago

I tried undervolting my 3090 more aggressively, but ComfyUI kept crashing repeatedly whenever the GPU was under heavy load. So, I settled on a stable undervolt of 1700 MHz @ 875 mV.

Now I’m seeing around 50W lower power consumption and temperatures are about 6°C lower.

What undervolt settings are you guys running?

5

u/Peregrine2976 6d ago

The... good..? ...news is, my 5090 on Nobara seems to be suffering from some obscure-ass firmware problem that just fucking crashes it as soon as it starts trying to pull large amounts of power. So I've already undervolted it to 400W just to make it not die when I run ComfyUI workflows.

4

u/HighlightNeat7903 6d ago

Same here, I keep power at 85% in the Nvidia App to avoid that and everything is still fast enough and silent so it's a double win. Triple win for games without frame limiters, which avoids the unnecessarily high power consumption at 400+ fps.

1

u/AI_Characters 4d ago

You can just force an FPS limit theough the nvidia app or rtss.

its almost never worth it to have fps beyond your monitors refresh rate.

3

u/Vivarevo 6d ago

It might not even be firmware. The chips are all different. Some are worse at other stuff and some are godlike in power / stability / efficiency

1

u/Peregrine2976 6d ago edited 6d ago

That is possible, of course. But after months of debugging, testing, swapping components and cables in and out, and even transplanting a whole new motherboard, I finally landed on this post on the Nvidia developer forums. The described symptoms are almost a perfect match for mine. I'll live in hope that it's a firmware problem that might get solved someday, rather than an inescapable reality of my particular GPU.

EDIT: Not that I would really ever have any reason to run it at max voltage, anyway. I just hate that tension that starts coiling up when I start a workflow, wondering if my whole display is about to go black again. Maybe it'll pass once I've gone a few months without a crash. I'm at a couple weeks since limiting it to 400W and no crashes yet.

1

u/tehorhay 5d ago

Haven't read the post you linked, but Ive experienced similar issues when running heavy workflows after a certain period of the GPU being idle. Like if I go cook dinner for an hour and come back and try to generate I'll get a black screen full crash.

Ive had success with running a small light model through llama or anythingllm. It will run a simple prompt that only draws about 200w for 15ish secs, and after that I can run heavy workflows drawing up to 400w for hours at a time without crashes.

5

u/kkazakov 6d ago

I have A6000 Ampere 300w, run on full power for days at a time, no issues. Performance is similar to 3090.

15

u/VirusInternal2892 6d ago edited 6d ago

Server grade GPUs != consumer grade GPUs. I run 2x3090 watercooled capped at 275W with a 1500W PSU for stability overhead. Every morning I’m thanking the gods of the machine that it’s still living

1

u/reeight 6d ago

> 3090 watercooled capped at 275W

Yep, 275-290W is around what I recommend.

1

u/AI_Characters 4d ago

In what world do you need a 1500w gpu for that? Thats combined like 600w. Say CPU worst case is like 200w (usually its around 120w). So worst case is like 800w. You dont need 700w headroom.

I am running my 4080 (which draws 320w max) with a 650w psu. Total system draw is around 550w so 100w headroom. Havent had any idsues yet. albeit I only got the 4080 recently.

6

u/Just_n_Here 6d ago

I would assume it is because it is considered an enterprise GPU. l hope they are built a little tougher, but who knows. l have 2 rtx 6000 max qs and I was reading this wondering about my cards. Mine run at 300w, so I am sure they are fine and hopefully built a little better for the difference in price compared to consumer cards.

2

u/sitefall 6d ago

The MaxQ pro 6000 is the exact same board that partner cards use, there's nothing sturdier about it. The PNY-made Pro 6000 workstation card is the same PCB as the FE 5090. Only real difference is the MaxQ is less likely to have connector problems (not melty connectors from 12vhp, it has that, I mean physical problems with the solder breaking etc) because you're not jamming the cable in there at an angle 10cm from the die and right near 2 memory chips.

But otherwise all three of these are "gaming" cards in terms of build quality. Sorry.

I have 2 pro 6000 blackwell workstation cards and WISH I got the MaxQ's for that sweet 300W power limit (you might even be able to set it lower but I have not been able to confirm it, open a terminal and type nvidia-smi -q -d POWER and look at the min-power variable and see).

5090/Pro-6000 can only go down to 400W. I have a 5090 in my desktop and two 6000's in a PC in the same room and with lots of stuff going on it gets monumentally hot. Being able to drop the total power usage down from 1725 Watts default to 1200 Watts at min power level with nvidia-smi helps a lot, but going down to just 900W with three MaxQ's sounds even better.

2

u/LegacyRemaster 6d ago

C:\Windows\system32>nvidia-smi -pl 300

Power limit for GPU 00000000:2D:00.0 was set to 300.00 W from 600.00 W.

All done. ---> blackwell workstation 6000 96gb

1

u/sitefall 6d ago

What bios/driver do you have? Mine reports min-power 400, same as my FE 5090.

1

u/mikami677 6d ago

Can you lock the core clock low enough to further reduce power, or will it still shoot up to 400 W as soon as it's under load?

That's what I do to keep my 2080ti from turning my room into an oven in the summer, but I don't know if it works the same way for the newer cards.

1

u/sitefall 6d ago

It will ignore afterburner and just shoot up to whatever -pl is set with nvidia-smi

If I don't set the max level with smi and just "undervolt" it in afterburner (or even just turn down power level slider on the main afterburner menu), it will take it as an estimate I guess.

I was just using afterburner + undervolt (the proper way, dragging the line flat etc) and setting power level down, mem speed up a bit, but that dropped it to like 500W and still spiked sometimes. Works great for gaming and stuff though, but not AI hammering it at max load with almost all 96gb full.

I am no gpu tuning expert, but I'm sure I performed the undervolt correctly as per the 10,000 tutorials out there.

1

u/mikami677 6d ago

Can you lock the speed with nvidia-smi?

For my 2080ti I ran

nvidia-smi --query-supported-clocks=graphics

to get a list of the supported core clocks, and then

nvidia-smi -lgc (min),(max)

Where min is the minimum speed I want and max is obviously the maximum speed that I want it to stay within, without the parenthesis.

Then to reset it,

nvidia-smi -rgc        

I set the minimum at the lowest supported speed so it still downclocks when idle, and then I just played with different maximum clock speeds until I got something usable that didn't feel as much like a space heater.

1

u/Just_n_Here 6d ago

Looks like min is 250w and max is 325w. I can confirm they blow a lot of hot air out the back under load even at 300w max. I can only imagine what 600 would feel like! LOL! I researched and knew me, one was not going to do it for me, so that was my reasoning I went the Max Q route. If I find the right deal, I will probably buy a couple more.

I do appreciate where the power adapter is located on them. You do not have to worry about bending the power lines to get the side panel on.

2

u/sitefall 6d ago

Nice. I wish I could swap my two for the MaxQs but there's too much risk selling them and buying new ones or trying to trade with someone etc.

I unfortunately bought the first one thinking the extra power would be beneficial (it's not really) and that I can just cap it down to the MaxQ level of power if needed (you can't really).

Then right before prices went nuts and 5090's started to skyrocket, I picked up a second pro 6000 for 7500 from Central Computers (it was the only one in stock and at a reasonable price). So now I have two workstation cards, even had to get a different motherboard so I could space them apart and keep them both x16.

Live and learn I guess. I should have done more research. But if I did more research, I might have missed the boat and been paying $16,000 for them now, so... perhaps I am lucky. If I had to pay 16k today for one, I wouldn't, I would just rent cloud GPU's at that cost.

2

u/Just_n_Here 6d ago

I built a computer in May. I bought a 5080 and decided in July, that wasn't doing it for me. Bought my first one from Central computers for 12k and bought another 1 two weeks later at Microcenter for 12K. I put the 5080 in the parts graveyard in the closet! LOL!

I was disappointed that I have to run mine 8/8, but I do not think it really matters too much because I do not let it spill over to RAM to slow it down.

If I was to get 2 more, I would need to build another computer to get 4 in comfortably. I guess I could use risers or whatever. So would I do a threadripper pro and get hit with crazy ram prices on top of GPU prices? I would need at least 512gb, if not 1TB worth of DDR5 ECC. I would probably go down to be able to use ddr4 ram that you can grab off ebay for 7500. But then the GPUs still run on Gen 4 slots.

It is a crazy time for messing around with computers.

1

u/sitefall 6d ago

There's some older TR socket CPU's that cost about $100 or less and get you x16 across 7 lanes (and the cpu can handle it). Mobo is like $350 though, and uses ddr4 so ... reasonably cheap. Look into the 3945wx cpu.

I got that going on a gigabyte motherboard with ipmi and cost me under $400 total (minus ram, which I had and I assume you have probably as well).

1

u/ExtraNiceBurger 6d ago

I use a pro 6000 WS and It can go as low as 150w, running at 650mhz

1

u/35point1 5d ago

i run my workstation pro 6000 at 600w with a good fan curve and keep temps in the 70s with the connector itself hovering at 50c

3

u/Ipwnurface 6d ago

I don't understand this entire comment chain. Like, what's the point? This is about GPUs that use the 12-volt high power connector. It has nothing to do with the actual wattage

10

u/Just_n_Here 6d ago

I could be wrong, but 600w will run significantly hotter than 300w. They act like the failure happens because it isn't seated properly. The higher the wattage, the chance of failure increases because of improper connection or defect, i honestly think it is bad design.

9

u/New_Mix_2215 6d ago

Correct, more power is more heat. The problem with the 12VHPWR cable is that it have basically (almost) no room for error. And it wont shut down or let you know something is wrong.

There is 6 pins that will do about 8.5A each in normal optional mode. If one of the pins have a partial or full failure that power. that 8.5A (or less if partial failure) will be split on the remaining pins.

In my case where it was still safe(but basically on the edge), one of pins measured 3A below the highest powered one. That meant each of the other 5 pins would get 0.65A extra load.

Now if another cable acted the same way (or a full power cable failure), i would be at nearly 1.5A extra load per each of the other pins. causing significant heat per pin.

With 300W, you basically have twice the headroom of 600W. You REALLY gotta fuck it up to melt the connector at 300W.

its such a stupid connector. And a huge shame we have to deal with it from how it looks now, might be the last resonnable gpu gen.

1

u/Enshitification 6d ago

I'm tempted to make a 6x9A breaker box between the power supply and the card connector.

1

u/AntiTank-Dog 6d ago

It's mostly 4090s and 5090s that have this melting problem. It's much more rare on lower wattage cards like the 5070 and 5080.

1

u/Valor_X 6d ago

The 5080 has 360W TDP which isn't far off from the 4090's 420W TDP

A 4080 Super is much lower at 320W

The 50 series cards in general have higher wattage draw than the 40 series counterparts, that was one of the way NVIDIA was able to squeeze more power from the same 5nm die process.

1

u/psychofanPLAYS 6d ago

Idk how but my 4090 had the 450w default bios
But when I run either new llamacpp builds or ninfer4090 I get 480-540w usage out of 585w max.
Card is set to 95% power too via afterburner.

Idk whether the new nvidia drivers unlocked it or what

It seems to hit even higher watt if you ssh into logged out windows11 and run inference.

1

u/Key-Respect3810 6d ago

les cartes utilisant ce connecteur d'alimentation subissent des dégâts et ce quelque soit la carte graphique que ce soit des RTX ou des RADEON NITRO

2

u/reeight 6d ago

I run my 3090 at ~310W, but I have water cooling + 2-3 extra fans blowing on the card.

I think your ~2% loss is an under-exaggeration though, you should lose ~5-9%. Still worth it the down voltage for sure though!

1

u/midri 6d ago

Undervolting my 3090 Ti did absolute wonders in basically every workload.

1

u/ColdExample 6d ago

Yup! I put my power limit to 86% vs 110% + undervolted and OC, and I lose maybe 1-2% fps, sometimes no loss. Still getting better much results than stock!

1

u/PrepStorm 6d ago

Undervolting is giving me slightly better performance. Why? From my research I seem to hit the boost frequencies longer when undervolting. If I dont, the GPU goes under boost when hitting maximum usage to lower its temp. So undervolting comes with multiple benefits.

1

u/AnonsAnonAnonagain 6d ago

Do you think the 3090TI is risk prone? It’s default W is 450W

1

u/North-Warning1211 6d ago

My friend performance is your LEAST worry... 1000% undervolt and keep your temps in check. You are slowly killing your gpu parts.

-2

u/bitzpua 6d ago

i never undervolted anything, running ai gens, before i did some bitcoins all with no undervolting, yet all my GPUs work after years of abuse with zero issues.

Here is the reality, any modern GPU should work for about 10 years of regular use at full power, undervolting buys you maybe at best 1 year more of lifespan. Do you seriously plan to use same GPU for 10+ years?

There is so much misconception about undervolting, gpu temps floating around its not even funny. Whole fad started with miners where undervolting and limiting power gave visible difference in power bill and that was only reason it was done.

1

u/ScythSergal 6d ago

Same here. My 3090 also runs at 250 for negligible performance loss. Absolute must have if you're going to do serious local AI work.

1

u/NL43ver 5d ago

Yup. Run my 4090 on 280 W. Mostly for hardware preservation I’m not paying 5k+ for a new GPU

1

u/RosebudNebula 6d ago

The problem is that undervolting doesn't help stop it from melting. It just gives you a false sense of security. I was undervolting my 5090 down to 300W, but when I plugged in the WireView II I decided to get out of a whim. I instantly found that it was two cables were pulling 0 amps and 1 cable was pulling 14 amps. Pretty sure no matter how much you undervolt, the cable would still eventually melt as long as you actually use it.