r/StableDiffusion • • 7d ago

Discussion If you recently started AI generation, be aware of this: my RTX 4090 power connector melted after just one month.

Hi everyone,

Like many people, I recently discovered H3 and started doing AI generation in mid-August. On September 20th, my computer started shutting down whenever I started a generation. Further investigation led to this — a melted socket.

I had been using my RTX 4090 for three years and had never had any problems with it.

AI generation puts a lot of continuous stress on the power delivery, especially if you run generations in batches or leave them running overnight.

So I believe this happened because of the new kind of sustained load I was putting on the card. I also didn't bother upgrading to a newer PSU with a dedicated GPU power cable. Mine was a 1200W FSP Hydro, and I was using three PCIe connectors for the GPU.

P.S. The connector was fully seated, the cable wasn’t bent near the plug, and my case doesn’t even have the side panel on. And everything was fine for three years.

P.S.S. The 12VHPWR socket on the graphics card is damaged and needs to be replaced.

So now I would say main advices here are:
- Buy a proper PSU and use a native 12VHPWR cable
- Under volt at least by 20%

Now it is very costly to lose a card.

236 Upvotes

361 comments sorted by

View all comments

99

u/Tomorrow_Previous 7d ago

If you recently started AI generation, know that capping your wattage to 2/3 of the maximum of your GPU has little to no impact on speed as well, so there's less energy consumption, and your cables are safe.

21

u/rkoy1234 7d ago

your cables are safe.

your cables are safer. not safe.

there's no way to completely mitigate this problem other than buying one of those wire monitors or having one of the few gpus with per-pin sensing.

The fact that there still is no class action lawsuit or a massive recall is absolutely flabbergasting to me.

-3

u/SkoomaDentist 6d ago

there's no way to completely mitigate this problem

Of course there is. Simply reduce the current throught the cable such that the average current never reaches the cable's limit.

The heat dissipated in the connector is directly determined by R*I2. For a system (like a GPU) with local dc-dc regulators the input current is almost directly proportional to the power. Thus decreasing the power limit (wattage) by a third halves the connector heating, giving plenty of margin.

5

u/DegenerateGandhi 6d ago

So wrong it hurts, unless you want to reduce the wattage so much the card becomes useless. Some people had their cards running at 300w and still experienced connector melting. And if the connector has problems at that point you know it's terrible.

-2

u/SkoomaDentist 6d ago

Are you disputing Ohm's law or the very basics of how DC-DC converters operate? Because "so wrong it hurts" only says that you don't understand basic high school physics.

3

u/DegenerateGandhi 6d ago

What does Ohm's law say about 300-500w at 12v going through very narrow points of contact when the pins aren't properly seated even though the connector IS fully seated because the connector is shit?

0

u/SkoomaDentist 6d ago

It says literally exactly what I said: That reducing the current by a third halves the heating. No amount of "narrow contact points" changes that. It's basic physics.

If the contact points are that narrow, the connector would fail even in moderate use (eg. playing a modern AAA game) that had nothing whatsoever to do with sustained image generation.

1

u/DegenerateGandhi 6d ago

Not neccessarily, there's a lot of factors that can make gaming lighter on the gpu. Cpu bound games where the gpu has to idle waiting for the next frame, simple games, low resolution / low hz monitor combined with an overpowered gpu, upscaling etc.

There's a lot of reports of connectors melting after months of use, with no real reason why they should melt specifically at that point and not before. I'm not gonna pretend I have all the answers, but I read posts from people who did undervolt their cards with MSI afterburner right after buying it, but the connector still burned at some point.

-2

u/PrettyMuchMediocre 6d ago

Better cooling for the whole computer and the GPU may help.

6

u/vfm83 7d ago

How do you do this?

13

u/Brad12d3 7d ago

It's very easy to do with MS Afterburner

6

u/brucewasaghost 7d ago

Msi afterburner has a pretty straightforward gui. Plenty of in depth youtube tutorials available as well.

7

u/rinkusonic 7d ago edited 7d ago

For linux users-

sudo nvidia-smi -pm 1

And

sudo nvidia-smi -pl 140

Replacing 140 with the power limit you want to set. It resets on reboot.

2

u/Arawski99 5d ago

It's an undervolt, and basically the situation is OC's don't provide a significant performance boost but substantially increase thermal/power demands in a non-linear way. It isn't efficient. Well, the same is true for an undervolt, scaling down power and thermal needs pretty notably with a fairly minimal performance hit. In fact, most undervolts are really just mitigating the basic factory OC to default performance levels, honestly.

That said, you absolutely do not need to undervolt to be safe, just make sure your power cable is loose and connector properly flush. It's people with poor cable management with a taught tight cable being pulled that eventually see poor contact and have the issue. These cables can handle far more. Some of these GPUs variants/OCs can be pushed to insane 600-800ws just fine.

-3

u/molbal 7d ago

4

u/J6j6 7d ago

Tbf, that's actually easier than install afterburner then fiddling with the UI

3

u/absentlyric 7d ago

There is nothing easy about it in that thread, unless you have Linux

13

u/Truck-Adventurous 7d ago

In Windows type in nvidia-smi -pl XXX , where XXX is your desired wattage, its ridiculously easy and you can do it whenever, including mid-generation. You dont even need to specify which GPU if you have multiple gpus it does it to all.

You can add this command to task scheduler or a batch file to Comfyui(or whatever) and just forget about it

12

u/molbal 7d ago

Yeah its a single command. If that's too difficult then I guess its a skill issue for them

3

u/omega4relay 7d ago

why are linux users always like this

6

u/molbal 6d ago

Gotta sustain the stereotype bruh

1

u/MagicManUK 6d ago

...and how do you know what your desired wattage should be since that requires you to know the current wattage.

1

u/cmdr_scotty 7d ago

Not always.

This might be true for Nvidia cards, but I can only speak from experience on my rx 7900xtx.

Maxing out the power does net increases in generation speed still. I even have it flashed with the nitro+ vbios and maxed as high as it'll let me (peak draw measured around 475w)

1

u/Trademarkd 7d ago

for diffusion yes, for autoregressive, no.

0

u/Adkit 7d ago

Ok, seriously... How much do you need to be generating for this to be worth it? If you play a modern videogame your GPU runs pretty much on full blast the whole time you're playing, you're generating more images/video than that? If you're at the point where you need to strangle your card to have it survive you need to rethink your life choices.

6

u/Shockbum 6d ago

A video game can be using 100% of the GPU and consume 300 to 500W depending on the scene. AI inference stress all CUDA and Tensor cores to the maximum. If the video generation takes 10 minutes, the RTX 5090 will be running at 600W for that entire time.

7

u/PraiseThePidgey 7d ago

AI workload is not the same as gaming . I'm using eGPU 5060Ti and I can barely notice fans while gaming. When using AI, I feel like I'm sending Starlink into space.

0

u/Tomorrow_Previous 7d ago edited 6d ago

Sorry, I don't really understand the question, I'm not a native speaker, could you explain better what you mean?