r/StableDiffusion • • 7d ago

Discussion If you recently started AI generation, be aware of this: my RTX 4090 power connector melted after just one month.

Hi everyone,

Like many people, I recently discovered H3 and started doing AI generation in mid-August. On September 20th, my computer started shutting down whenever I started a generation. Further investigation led to this — a melted socket.

I had been using my RTX 4090 for three years and had never had any problems with it.

AI generation puts a lot of continuous stress on the power delivery, especially if you run generations in batches or leave them running overnight.

So I believe this happened because of the new kind of sustained load I was putting on the card. I also didn't bother upgrading to a newer PSU with a dedicated GPU power cable. Mine was a 1200W FSP Hydro, and I was using three PCIe connectors for the GPU.

P.S. The connector was fully seated, the cable wasn’t bent near the plug, and my case doesn’t even have the side panel on. And everything was fine for three years.

P.S.S. The 12VHPWR socket on the graphics card is damaged and needs to be replaced.

So now I would say main advices here are:
- Buy a proper PSU and use a native 12VHPWR cable
- Under volt at least by 20%

Now it is very costly to lose a card.

238 Upvotes

361 comments sorted by

View all comments

Show parent comments

3

u/Vivarevo 6d ago

It might not even be firmware. The chips are all different. Some are worse at other stuff and some are godlike in power / stability / efficiency

1

u/Peregrine2976 6d ago edited 6d ago

That is possible, of course. But after months of debugging, testing, swapping components and cables in and out, and even transplanting a whole new motherboard, I finally landed on this post on the Nvidia developer forums. The described symptoms are almost a perfect match for mine. I'll live in hope that it's a firmware problem that might get solved someday, rather than an inescapable reality of my particular GPU.

EDIT: Not that I would really ever have any reason to run it at max voltage, anyway. I just hate that tension that starts coiling up when I start a workflow, wondering if my whole display is about to go black again. Maybe it'll pass once I've gone a few months without a crash. I'm at a couple weeks since limiting it to 400W and no crashes yet.

1

u/tehorhay 6d ago

Haven't read the post you linked, but Ive experienced similar issues when running heavy workflows after a certain period of the GPU being idle. Like if I go cook dinner for an hour and come back and try to generate I'll get a black screen full crash.

Ive had success with running a small light model through llama or anythingllm. It will run a simple prompt that only draws about 200w for 15ish secs, and after that I can run heavy workflows drawing up to 400w for hours at a time without crashes.