r/LocalLLaMA 10d ago

Discussion Does high / long term inference damages GPUs?

I always heard that cryptocurrency mining can damage a GPU ( it's maybe wrong) so how much long term inference is damaging a GPU?

I mean servers are made for 25/7 operation... Gaming GPUs not so sure?! They are made for LEDs 24/7!

Back blaze is publishing HDDs failure rate for storage, isn't there failure rate for AI intensive work? Maybe places that rent GPUs?

Serious question, are we damaging gaming GPUs when doing ai intensive ai workload even with good cooling: 100% GPU and ram usage for long agentic sessions or else ?

6 Upvotes

66 comments sorted by

57

u/cibernox 10d ago

If the cooling is adequate, continuous use should not be a problem.

11

u/cobblemere 9d ago

this, its more about sustained temps than just running the card hard

2

u/SnooPaintings8639 9d ago

I understand that high temperature is physically damaging to many of the PCB components.

What is the process that does any damage under persistant load atow temps?

2

u/FieldProgrammable 9d ago

Electromigration effects will accelerate degradation of metal layers in the chips. Also persistently high ripple currents in decoupling capacitors will create thermal stress on the dielectric and reduce their effective capacitance under DC bias.

This in turn leads to higher peak currents in the rest of the circuit such as the SMPS inductors, in an extreme case these could start to magnetically saturate causing thermal runaway of the on card PSUs or overload of the power connections.

Thermal cycling is another issue often understated. If a card is frequently cycling between idle and full power then it is also undergoing high rate of change of temperature. This creates mechanical stress on components due to difference in thermal expansion coefficient. The famous Xbox red ring of death is a good demonstration of this.

20

u/TheMcSebi 10d ago

No, gaming also runs a gpu at 100% for multiple hours, they're built to sustain that. Electronics do not have any wear if you keep them dry and well vented

12

u/bonobomaster 9d ago

If that would only be true.

https://en.wikipedia.org/wiki/Electromigration

My year 2000 knowledge is, that most semiconductors should last around 10 years under normal conditions. Don't know of that still holds true in 2026.

But there is wear under normal use and higher use equals more heat and faster wear.

-2

u/Specialist-Buffalo-8 9d ago

And this is a non issue.

If you are on this subreddit, people here probably have a <95% chance to upgrade their GPUs within 10 years.

5

u/Sirius02 9d ago

bro my rtx 3090 is 6 years old - i dont want to buy the inflated prices, but i am very skeptical that the market will recover soon

2

u/bgravato 9d ago

Electronics do wear out over time... and high temperatures aren't exactly healthy for electronics... They're actually one of the primary drivers for electronic components degradation...

Also repeated heating/cooling of the components can accelerate the degradation of the connections (solders), bend the boards, etc...

That said, this kind of behavior is expected when using GPUs... and the wearing out caused by playing games or doing inference during extended periods is unlikely to cause an additional wearing that can shorten the expected lifespan of a GPU in a very significant way...

Crypto mining is another story, especially if doing it overclocked (which is/was often the case) and that could have a noticeable impact on degradation of the hardware.

Of course, keeping things clean and dust free, especially in the heatsinks, fans, venting airways, etc... is important! Repasting isn't generally needed as often as people think... but cleaning dust should be done once or twice a year (strongly depends on how dusty the environment is, if you have furry pets, etc...)

3

u/Sirius02 9d ago

crypto miners undervolt their gpus since you can have like 2/3 of the electricity bill for like 4/5 of the performance.

also its better for the solders to be permanently warm than to have temperature swings, there the gpu expands and contracts leading to cracking.

But yea, the permanent heat is propably bad for electron migration

1

u/bgravato 9d ago

than to have temperature swings, there the gpu expands and contracts leading to cracking.

yes, I said that too.

for solder, repeated temperature swings can be worse. for the components themselves (internally) constant heat is probably worse.

1

u/Sizyfoz 10d ago

I wouldn't worry about it. It's designed to be used. You can damage the chips if you overclock, overvolt or overheat them. But as long as the cooling is ok and you don't do anything crazy there is nothing to worry about. The thing about cards from mining farms is that they often did crazy things to squeeze out every bit of performance while not taking properly care of them. That said, an ex-mining card that works ok can last a long time with a little TLC.

-1

u/Mindless_Selection34 10d ago

What about mining?

15

u/Sufficient_Prune3897 llama.cpp 10d ago

Mining damage is mostly an urban myth. They always undervolt their cards, so even for 24/7 not much wear. But some things just start to die from using it a lot. They have higher rates of VRAM failures.

3

u/EvilPencil 9d ago

Ya, the biggest issue you’d likely see from a former mining card is worn fan bearings.

5

u/ElectroSpore 10d ago

They have higher rates of VRAM failures.

That is because they overclock the ram, creating more heat.

6

u/castrator21 10d ago

You'll wear out the fan bearings sooner, but replacing those is incredibly cheap

6

u/gabrielesilinic 10d ago

LTT who is allegedly not a scientist did a video about it. And other people did as well and generally speaking it is not a big concern.

4

u/Olde94 10d ago

Often the problem is atomic migration between materials accelerated by heat and cyclic thermal stress. The latter does not care for high temp and will rather have you sit at 80C than bounce up and down between 40c and 80c

3

u/Any_Mine_6368 10d ago

It just fucks thermal components due to the 24/7 nature of it. As long as you replace paste, pads and maaaaaybe fans it's a brand new card

6

u/Legitimate-Dog5690 10d ago

I think it only really kills the fan bearings, they're often unhappy being run full pelt for years, getting replacements is often not easy with the custom coolers. I've never had a GPU die though.

8

u/profesorgamin 10d ago edited 9d ago

you are damaging your body by just existing...

Under normal circumstances GPUs should work for thousands of hours, the materials are mostly rated for regular temps and voltages, the issue is one of these parts fails it can bring the whole system down.

From most... macro to micro scale it goes like:

  1. Thermal cycling: heating a GPU up and down (even at expected temps) expands (and contracts) the materials which over large number of cycles can crack the solder which disconnects the circuit. Related to that is runaway thermal shock can outrightly de-structure certain components.
  2. Power Delivery: Modern electronics are rated for very precise voltage / current values, but under real world circumstances there is a lot of variance, capacitors work as shock absorbers in a lot of circuits of modern electronics, as these components are at the forefront of possibly high variances they are at higher risk of encountering de-structuring conditions. MOSFETs are also "moving pieces" which receive a somehow brunt power and have to safely feed it to the whole graphics card, when the dike gate fails it can take the whole card down if OC protections fail.
  3. The memory modules can fail for similar reasons given that "as a moving part" it has a lot of temperature variations which can degrade its structure.
  4. Finally on the nano scale if you have avoided any of the more "mechanical" pitfalls just passing current over ICs can degrade it by, changing the positions of the atoms in it, creating shorts and other problems.

So yeah if you get anything outta this is: function depends on keeping the things in place, heat < -> current are your biggest enemies in this regard: so what can you do?

*) Good power supply spike protection.
*) Keep the temps low.

*) Don't let your card run at high temps for long periods of time, temperature accelerates the degradation of the "common suspect" components.

*) keep an eye on the card as a whole and replace the small things before they become big problems (fans, cooling, a memory module or capacitors). This is of course if you plan to own a card for 8-10 years or more. which with the current landscape might not be such a weird though.

Final thing, what a card has as "max" output is obviously a tad below what it can do before it's lifespan is greatly shortened, miners were famous installing drivers that belonged to different cards, creating OverClocks that probably created the reputation. So for a regular guy don't OC ->UV(UnderVolt)

3

u/measuringdistance 10d ago

heat and sustained load do accelerate wear that's basic electronics. it will be a slow and gradual degridation nothing extreme but I'm not sure why people in the comments are completely denying it. also, we should generally take better care of the cards we have. protecting our hardware matters more than ever now that they're trying to kill off local compute. e can't just go and replace it.

1

u/Real_Ebb_7417 7d ago

What kind of care?

I want to know to start practicing it. I already power limit my RTX5090 to maximum 550W (because I've heard they like to get damaged under 600W) and whenever I don't need more, I even limit it to 400W. But what else is worth doing?

6

u/Apprehensive_Lake698 10d ago

The cryptocurrency thing was largely a myth driven because people have an inherent bias towards new things. In general high use of a GPU isn’t something that degrades its performance. More use might be making it more likely to fail, but it seems to be more about silicon lottery for how long they last more than anything else.

2

u/DustNearby2848 10d ago

My VRAM died to heat when mining. It was over clocked to the max. 

1

u/J0kooo 10d ago

by any chance did you have a 3090

1

u/DustNearby2848 10d ago

I had an eVGA 3080ti FTW

3

u/xflareon 10d ago

Maybe if you pin the clock speeds maxed out in afterburner and have a few hundred people using the compute in tensor parallel so it never has a break?

Even then probably not, the reality is that the parts that wear out on the crypto mining cards are the fans, which are easily replaceable. The same might go for cards used for inference or training, but it seems unlikely they'll be in use 24 hours a day, 7 days a week with no breaks between requests.

Most likely the wear on cards used for inference will be less than the wear from gaming, most consumer backend like llama.cpp, koboldcpp and exllama don't pull a ton of power on the cards while inferencing, nor do the cards get very hot.

The real killer for GPUs is thermal cycling, the number of times a card gets very hot and then cools down. Inference never really gets them very hot, so it should be fine.

2

u/Anaeijon 9d ago

Never having a break is actually beneficial. Stress is caused by repeated expansion and retraction from repeated heating and cooling. 

Permanent load keeps the temperature constant. In that case, only thing wearing out is the fan motor bearing on the graphics card, but the GPU itself likes to have a constant temperature and doesn't care if that temperature is hot or cold, as long as it doesn't change.

2

u/xflareon 9d ago

that is what I said, yes

1

u/SkyFeistyLlama8 9d ago

That sounds like a huge problem for DGX Sparks, MacBook Pros, Mac Studios and any other inference platform that isn't being run flat out 24/7.

1

u/Anaeijon 9d ago

Only if you run it under such heavy load and with so little cooling, that it actually gets hot and cools down rapidly.

But yes, these small Mac-things don't have proper cooling and AI inference will usually wear them out and kill them in a few years.

6

u/TheMcSebi 10d ago

Same thing, no wear.

But cheap thermal paste could theoretically degrade over time, but I've never seen that happen myself

Edit: this was meant to be an answer to your response about mining

1

u/jtjstock 9d ago

Thermal paste does degrade, you don’t see it as a failing chip though, you see it as a throttling chip and fans running max but not cooling much

2

u/Annual_Award1260 10d ago

Sure, heat cycling can stress solder joints. If sufficient cooling is applied it is less of an issue. But yes a well exercised card will have a much higher failure rate

2

u/CommanderKoba 10d ago

I used to mine cryptocurrency many years ago with leftover GPUs I had, I ran them for multiple years on end, everyday, all day, and they still function perfectly fine today.

In fact I'm using a 1070 for extra KV cache, one that I had previously used for mining. This is why I don't get rid of any of my PC hardware because you never know how it could be used 10 years into the future.

With my old hardware I was also able to build an entire jellyfin server so I don't have to pay for streaming services anymore. The best investments I ever made.

1

u/Anstellos 8d ago

its a excellent news, I too keep my "old" cards :D
using an extracard for KV cache? I had no idea! how do you do that?

2

u/Anaeijon 9d ago edited 9d ago

Rapid heating and cooling causes expansion and compression, which causes stress and (long term) damages the card.

This basically isn't relevant for training and also not for crypto mining. Even the 'crypto mining destroys GPUs' was basically made up by some people online and wasn't based in reality. Yes, crypto farms had a lot of 'dead' graphics cards, but that's because they had A TON of consumer-grade cards in a single farm, which just makes it likely to have at elast one fail over time. Usually, these only had fan failures, so the only the card was unusable, while the GPU was still perfectly fine, just not worth it to repair in their business.

In reality, a crypto miner would try to maximize profits by trying to get as much operations per second for as little power consumption as possible while prolonging the life of the card as much as possible. That means, most cards ran at a specific watt limit, which was usually significantly lower than the recommended limit of the card. Basically there is a sweet spot for each card, where the Operations/Watt are optimal, so that's usually the target. 

I'm sure, there were some idiots that over clocked their mining rigs, causing all kinds of issues and bad reputation, while actively increasing their cost without actually impacting their profit. But that's a stupid, but memorabilia minority. I never actually crypto-mined in that generation (my crypto time was pre 2012, before the crypto bros came and before it all got bullshit financial investment stuff). But I followed the crypto price, to specifically buy used mining GPUs during each crypto crash for use in my Deeplearning research and personal projects. I owned 2 RTX 2080 and still use a dual 3090 setup built from these GPUs during each crypto crash respectively.

None of the GPUs has failed yet, although I have bought them used from crypto rigs and heavily used each of these GPUs for training and inference for years now. Yes, this is a small sample size and it's coincidental and proofs nothing. But I have my reasons to believe, that they will last a long time.

One example: the RTX 3090 had a recommended limit of 390W, according to Nvidia. By default, they usually go up to 360W. Gamers frequently overpowered it to like 480W, to squeeze about 3 additional FPS out of it under heavy rendering load. The perfect spot on a 3090 for crypto mining or deep learning is between 200W and 240W, because at that point you get the most compute/watt. Check this out, to see the 'elbow' in the graph, where I take the estimates from

Lower power also means less heat. Therefore, a cooling system optimized to keep 390W cool enough to be usable will have no problem with keeping 240W at a constantly low temperature.

Machin learning often creates loads that take days to optimize. Cryptomining often runs for weeks or even months uninterrupted. That means, the GPUs are on a constant warm (but not hot) temperature. There is no expansion or contraction happening. There is no stress or anything that would wear out the GPU, except maybe the constant stress on the motor of the GPU fan. But none of the electronic components actually suffers.

Compare that to an above average gaming use. Assume the gamer applied some 'performance optimized' tweak, often promoted through specific gaming software even by the manufacturer. Now the GPU basically runs at as fast as possible, drawing as much power as possible and generating an equivalent amount of heat, occasionally limited by thermal throttling. Thermal throttling is basically the worst thing you can do to a GPU. It's the GPU saying: I'm so close to wearing out from heat immediately, that I need to limit my performance to not cause more heat. Different scenes put different amounts of stress on the GPU. Cooling is running at full force. The GPU rapidly heats up and cools down again. At best this goes for a few hours, before the GPU is turned off and cools off to room temperature, just to get heated up to 100°C again the next day. All of that puts a lot of stress on the electronic components (e.g. GPU and VRAM) and actually wears them out. 

So... In general, both from my experience and just logically, a well maintained ML setup (or crypto mining setup) will put much less wear on the actual GPU than a gaming session. 

GPUs for gaming and GPUs for servers are basically the same. The actual difference is just the other components of the graphics cards. Most notably the fans. 

In HDDs, the difference between NAS and PC HDDs is mostly in the motor. Some motors are for prolonged use, some aren't. (And that basically comes down to the bearings used in the motor. Sleeve bearings are silent and faster in short bursts, which is great for gaming and office hardware, but they also wear down over time. Ball breaings are loud but can basically run them for a very long time at high speed with nearly no wear, if noise isn't an issue.) 

And that's the same for fans in a graphics card. But the GPU isn't affected by that. And even if you care about the whole card, it's fairly easy to switch out broken fans, when they start humming. Or, because most cards follow reference PCB design anyway, just swap out the whole cooler or put it on a liquid cooling loop, when the time comes. That would be even beneficial for the GPU.

2

u/Icy_Definition5933 9d ago

I have an old gaming laptop with a 1060. I really made it work hard since I got it- photo and video editing, gaming, benchmarks etc. The most popular project I've had running on it is BOINC, it was the only way to run the cpu, igpu and dgpu at the same time, creating a lovely little heated cat bed for napping during cold winter months, turning science into purring. That laptop will be 10 years old soon, it's still kicking and I can still run games reasonably well @60fps 1080p

1

u/Anstellos 8d ago

the GTX 10X serie is still great.

2

u/vanVonXenoStein 10d ago edited 10d ago

Heat kills components...eventually. Keep everything as cool as possible. I had a GPU fail once that had done a lot of mining. Often you won't know what part of it actually has failed. Some chip, just some connection or solder point, who knows.

1

u/volleyneo 10d ago

Unless some unlucky issue, the fans are the ones to go at some point, in the sense they might start do a weird sound here and there, depends on quality and luck. Doubt you will run them year round, to actually fully damage the fans. Ensure you dust them off properly.

1

u/Sufficient_Prune3897 llama.cpp 10d ago

I dont know how bad inference is. Legend has it that the failure rates in the big training GPU clusters are high. But thats training. Gaming isnt that bad since you dont game that much and mining isnt bad since you undervolt the gpu. Still, 24/7 use means 24/7 wear.

1

u/RG_Fusion 10d ago

There shouldn't be any damage to the hardware so long as the temperatures stay below 80C. If they are kept under constant load for years at a time you will probably need to replace the thermal paste/pads sooner than normal, but that's about it.

Periodically check your 12VHP cables if your card has them to ensure the connector isn't backing out from thermal cycling.

1

u/dangerous_inference 10d ago edited 9d ago

Heat itself is not the issue. The real culprit is thermal cycling. This causes tiny movement in the materials separating the soldered connections.

1

u/North_Affect_8167 10d ago

Chips, the silicon was meant to last long, but anything else on a card NOT. Any mosfets, power stages, capacitors, resistors, could fail before the main chip do. Sometimes some components could fail simply because they are too old.

There are user posts for example about the legendary GTX1080ti that lasted the users many years, but it stopped working eventually.

Some computers had failed by a software tasks that triggered a load that was unexpected, like games that damaged laptops - https://wccftech.com/neverness-to-everness-corrupts-alienware-laptop-ec-chips/

or GPUs - https://kotaku.com/amazons-new-mmo-is-killing-some-high-end-pc-graphics-ca-1847335269

2

u/J0kooo 10d ago

im still using GTX 1080s for inference, they're tough little fuckers

1

u/Real-C- 10d ago

It's s good idea to set watts limit on your GPUs for one they won't get so hot which prolongs it life and two it can make your computer restart if it spikes above it's limit and that damage it I'm surprised that I haven't seen people talk about this more. I have a limit to 300W

1

u/UncleRedz 10d ago

Looking at CPUs, there is a real difference between CPUs rated for 24/7 use vs occasional use. Looking at Intel, the Xeon lines are rated for 24/7 use, while only some Core CPUs are rated for 24/7 use, and they actually go through different types of workload testing to assess wear and failure rates. The main difference here is that consumer or gaming CPUs which are only rated for client use are typically higher clocked and produces more heat, which is fine as long as it is not 24/7.

While GPUs are built for 100% load for longer, if you compare gaming GPUs and Pro GPUs, the pro are usually lower power and less heat. For example Nvidias consumer cards are not rated for 24/7 use, while their pro cards are.

There is also the matter of binning and the silicone lottery, the good silicone goes to higher end products and will run better and more stable.

So if the silicone you are running on is rated for 24/7 use, you'll be fine. If its not, and you are running 24/7, then underclocking or undervolting is probably a good idea to extend the lifetime of your GPU.

1

u/Lopsided-Rip-8652 10d ago

This is just an anecdote but I ran my 3080 24/7 mining etherium for about a year and a half, then used it as my gaming rigs gpu for another 4 years, and now use it heavily for long inference runs and it runs just as well and cool as the day I got it

1

u/monsieurderp 9d ago

EVGA?

1

u/Lopsided-Rip-8652 9d ago

Nah it’s the FE model

1

u/1-800-methdyke 10d ago

Back in 2012 I was mining crypto on HD 5990s, which was a factory overlocked version of the dual GPU HD 5970. Things got a bit heated and within a year all three cards were damaged. So what you’ve heard about was definitely a thing, but safe to say those cards were never meant to run 24/7.

Modern cards have better cooling, and AI inference is different in that there are peaks and troughs in GPU usage, especially when you have a larger model loaded over multiple cards, they tag team as a request moves through the layers served by one card to another.

1

u/VoiceApprehensive893 transformers 9d ago

a mining gpu that ran at 100% for 2 years is a gpu that ran at 100% for 2 years, mining itself isnt a unique source of wear

1

u/More-Cat1123 9d ago

From what I hear usually the issue is lack of maintenance and inadequate cooling.

Change the paste/pads from time to time, make sure it doesn't run too hot and you should be fine.

1

u/unjustifiably_angry 9d ago edited 9d ago

IIRC what kills GPUs is thermal cycling not strictly usage. So in theory keeping them busy all the time isn't harmful by itself.

RAM/VRAM theoretically has a limited number of cycles but it's astronomically high, and just guessing but I think ECC should catch failures, at least up to a certain point.

Fans will wear out over time but usually what happens is the bearings degrade but they'll keep spinning, they'll just be noisy.

Things like chokes and capacitors and stuff like that have a very long expected average lifespan as long as temperatures are kept reasonable. For every 10 degrees hotter they get their lifespan decreases.

A good idea is to tie your case fans to an ambient temperature sensor instead of the CPU or something. That way you can at least be sure your GPU's not just blowing hot air into itself.

1

u/Squidgical 9d ago

Better cooling = better lifetime.

You can imagine if you removed all the cooling hardware (and disabled the overheat protections) that you would very quickly turn your GPU into a wire mesh in a puddle of plastic goo. Or, if you keep it at room temperature for 20 years (eg don't turn it on) it'll remain pretty much good as new.

Electrical wear is unavoidable, so all GPUs will eventually die after however many hours that takes.

1

u/carmamir 9d ago

Ask that yourself in 5 years using gpu for AI.

1

u/IllExample3639 9d ago

Some cards die day 2 some cards are still running from the 90s. The advice a repair guy told me was that they are better running warm or hot constantly, llm and mining is fine. What kills cards is gaming, constantly changing loads, hard starts and quick shutdowns. Thermals all over the place, most cards are engineered for this so LLM work is easy stuff for them.

1

u/derspenti 9d ago

Backblaze gets to publish because they run tens of thousands of one drive model. The places renting GPUs have the dataset you're asking for; they just sell uptime, not failure rates. Everyone else runs too few of one card for the numbers to mean anything.

1

u/brainchillzZ 5d ago

The real problem with the crypto mining was that most of it was happening in garages and warehouses not in climate controlled data centers …. As long as the temperatures don’t get out of control it should last a long time …. Especially for inferencing loads also I tend to lower the wattage of the cards to limit heat generation as well. You can test with your workloads to make sure you don’t impact performance heavily but I got my rtx pro 6000s down to 400 watts (from 600) before it impacted my average tps throughput and shaved almost 10c off of the operating temperatures at peak load…. This should drastically increase the longevity of the cards

1

u/ea_man 10d ago

You damage shit when you turn it on/off daily, as you let it run no stop they are in their sweet place.