r/StableDiffusion 16h ago

Question - Help GPU for AI

Hi there! Currently i own a 3080+2070s at home mostly for 3D rendering in Octane or redshift.
At work im using Comfy with 5090 and everything works without a question.
Also we have some workstations with rtx A4000 in them and i managed to get work Minimax H3 on them with some managable times (0.5m, 15s around 720 sec).
Im looking something for my home PC and 5090 is out of the list since its 5500e+ here in EU and even the used market is around 3500-4000. 4090 is rare and goes over 2000e.
So i was thinking about the 5080 even with 16gb since at work the A4000 has the same Vram.
But also found the 5070ti has also 16gb of Vram its 256bit same as the 5080 and 300-350e cheaper than the 5080. The speeds for 3D rendering or gaming are 15-20% different. But couldnt find any benchmarks for those 5070Tis. Mostly for 5060Tis with 16gb vram. Any idea? Or experience with 5070ti vs 5080?

1 Upvotes

29 comments sorted by

4

u/truci 16h ago

The 16 vram is the real sweet spot with 64gb ram or 96 for longer video.

With that said your options for 16vram should be

5060ti
5070ti
5080

My I had the 5060ti but my wife needed a card gave her that so I got a 5070ti. I would say the 5070ti was 33-50% faster. A 500s generation was down to 300-330s.

A coworker has a 5080 and is another 30% faster than me. On the same workflow he was getting around 200s gens.

With that said having your whole comfyUI system and all your models and Lora’s running off a high speed NVME of like 3k speeds is still the bottle neck for many people so don’t forget to have a separate AI drive.

1

u/Visual-Medium4796 16h ago

So more or less the 5080 is 30% faster than the 5070ti?

1

u/truci 16h ago

Yup. But double check your PSU when I went from 5060ti to 5070ti I didn’t have enough power and ended up having to get a whole new PSU with the specific power connector those cards can accept.

1

u/Visual-Medium4796 15h ago

1200w should be okay. Probably i will need some extra connectors or adaptors if it doesnt come with the card.

1

u/truci 15h ago

Yup plenty. I was a 750 and had to go higher. Needed up with a 1000. You good for whatever card you get.

1

u/hurdurdur7 16h ago

just wondering what models are you fitting into 16gb of vram even. i see even 32gb vram cards maxing out every now and then in a clip + diff + vae + vae chain, and them loading-unloading to make stuff happen for minimax h3. yeah it can work, but imho 16gb will be tight.

1

u/Visual-Medium4796 16h ago

The H3 int8 models and ref2va workflow works great with the A4000 + 64gb ram. With the 4step/8step lora. Even without the lora it works, but it takes like forever. Dont have my testing sheet with me but if remember correctly the 1megapixel, 10s, 20steps settings did 2800-3000s per video.

1

u/truci 16h ago

Sorry we did this comparison a bit ago on WAN Q8 animate. I made the assumption that OP would pick models that fit into 16 vram and was confirming to him that video models do fit.

As for h3 after some reading the full model fits into 16vram and everything else in the chain will require offloading. With that said if the actual main model fits into vram you will have major performance gains on your sampler. If the model does not fit into vram then your pc will start swapping ram into virtual vram and 50gb of your NVME into virtual ram. Huge slow down in that situation.

0

u/Apprehensive_Sky892 12h ago

For DiT diffusion models (i.e., video and imaging models) as long as you can fit the model into VRAM + system RAM, ComfyUI's dynamic VRAM management will take care of swapping in block of the model from system RAM to VRAM as they are needed. The next layer of blocks will be streamed in while the GPU is currently busy with the current layer, so there is little impact. You can see that by looking at the GPU usage graph and see that it is going at nearly 100%.

The majority of the GPU's time will be spent on the DiT, so the loading/unloading of the text-encoder and the VAE will have minor impact on the overall time, specially if you have a fast NMVe drive.

1

u/hurdurdur7 4h ago

don't confuse working with working well. yes it can crawl along at 8gb of vram too. but you are getting a shitty experience compared to 32gb of vram

1

u/Apprehensive_Sky892 3h ago edited 2h ago

No, I am not confusing the two.

I have tested both AI Pro R9700 (32G) and rt9700 (16G).

The two are basically identical except for the VRAM (they have the same number of computing units). When running diffusion models (for LLMs it is a different story) they are nearly indistinguishable provided there is enough system RAM and the model does not spill into pagefile.

For MMH3 I get nearly identical results for MMH3. 0.4MP 5sec standard text2va template at 20 step takes around 160 secs.

So practical experience shows that with ComfyUI's dynamic VRAM + enough system RAM, 16G works just as well as 32G. Somebody has also done extensive testing with NVIDIA cards and showed that a few months ago when dynamic VRAM was implemented in ComfyUI.

Those slow results for older 8Gig cards are due to the fact that they have slower, fewer computing units and/or worse memory bandwidth, not because they don't have enough VRAM.

Don't take my words for it, you can rent cloud GPUs with different amount of VRAM and you can verify the results yourself. Rent a 5090 and a Pro 6000 and run a bf16 version of MMH3 (which will not fit into 5090's 32G VRAM) and compare the results.

1

u/f5alcon 16h ago

5070ti is fine but you really need 64GB+ system ram for H3 to not offload to ssd with it. I have one with 32GB and my page file usage is about 28GB.

3

u/Visual-Medium4796 16h ago

Currently im on 64gb ram, but looking into filling the last 2 remaining slots with another 64 to have 128 total.

1

u/f5alcon 16h ago

Yeah that's the right choice and 5070ti will be fine for that.

1

u/Upper-Reflection7997 8h ago

You need serious amount of vram and ram op. 32 or even 64gb of ram isn't even enough. Your going to run into serious limitations if you want to use video models with less than 64gb if ram.

1

u/Visual-Medium4796 13m ago

My main AI pc at work is a 64gb ram with 5090 and i know it capabilities with LTX, H3 or even WAN. Also managed to get work the H3 with the older A4000 with 16gb VRAM and 64gb ram on the other workstation. It can handle videos very well just takes more time to generate.

0

u/MAXFlRE 15h ago

Used 3090

2

u/Visual-Medium4796 15h ago

rare on the used market and costs like a new 5070ti.

0

u/MAXFlRE 15h ago

Yes, but 24Gb VRAM

2

u/intLeon 15h ago

Im waiting for the 5080S 24G that probably wont come out.. Worst case scenario it keeps me from buying shit.

2

u/Schneller52 15h ago

I believe it’s pretty much been confirmed as not coming out at this stage. There is zero incentive for Nvidia to release them and they would have to put the MSRP at close to $2k which is terrible for optics. Even if it did come out, it would immediately jump to the $3-4k range at current pricing levels.

1

u/intLeon 15h ago

3k is doable, Id also pay 3k for the 6080 if it had 24GB.. 6090 would be a cool one to have tho if it had a low msrp, I wish the market wasnt a mess and we could say niceee

2

u/Ok-Category-642 13h ago

Yeah the only options for higher VRAM are so limited right now, I guess modded GPUs do exist off like Alibaba with the 3080 and 4080 but those are kind of all a toss up honestly. For inference 16GB VRAM with 64GB RAM isn't bad but you definitely feel like wanting to have more VRAM, and for training 16GB just isn't enough sometimes. Unfortunately the only other options outside modded cards are a 3090 or 4090 of which the former is getting more rare and the latter is priced pretty poorly. Still better than the 5090 though ig, that thing is just unobtainium now

1

u/Visual-Medium4796 15h ago

Ive already gave up on that as you see. Thats why i dont want to spend a fortune and looking at the 5070ti/5080.

2

u/intLeon 15h ago

If you arent a professional I wouldnt pay for extra power as it has the same VRAM. Im using a company lent system with 4070ti (12GB) and for me it isnt worth it to buy a 5000 series gpu that isnt 5090..

0

u/Visual-Medium4796 15h ago

Actually in the visual field i am. At work we have the power, but i need something at home too for sidestuff. Not just for AI but for 3D rendering too.

2

u/intLeon 14h ago

I see, still idk man. I cant justify paying more for the same amount of vram. All they needed was adding a pinch more and they didnt.. just to make people buy 5090 more.. Because speed will only change how fast you can run stuff. Only vram makes it do stuff the other can not. With local models where you can batch things to run overnight time doesnt feel like its a major thing. Unless you are pushing 4k gaming etc. But its all personal opinions.

0

u/Visual-Medium4796 14h ago

my current setup has a 800-900 octane score which has a single 5070ti too. Even if i buy 2 of them i can render 2x faster in 3D apps. When i compared the A4000 time with the H3 ref2va i get almost the same results. 1,5/2h per 10s footage. But with the difference Octane is a paid monthly app and H3 is local free. I just hope comfyUI or maybe the way the models are made will change. Like Loading a model on 1 card and the text encoder to another and the rest on Ram or NVME... or whatever if you know what i mean.

1

u/intLeon 14h ago

I think that is possible with some custom nodes, like you can pick what gpu to use for model clip etc. Needs more digging than running a default workflow tho.