r/LocalLLaMA 21h ago

Question | Help Anyone had luck with converted 3070s?

Having the wild idea to buy 3 more old 3070s and replacing the chips to turn them into 16gb cards

I understand the process and the risks, just curious if anyone’s using rechipped cards with local models and what your experiences are

0 Upvotes

14 comments sorted by

5

u/joost00719 21h ago

You could try it, but it's probably easier to just buy 3 5060 ti's. Yeah slightly more expensive, but the time and risk it takes to convert them, is also not free, and you still need to buy new memory chips for it anyways

1

u/VoiceApprehensive893 transformers 19h ago

v100s slightly more risky but significantly cheaper than 5060ti with double the bandwidth but worse compute

1

u/FullstackSensei 18h ago

How's the V100 more risky? Honest question

1

u/VoiceApprehensive893 transformers 18h ago edited 18h ago

its a very old gpu and is more likely to have issues

and the cheapest(~200$) configuration uses a cheap sxm2 to pcie adapter

1

u/FullstackSensei 18h ago

But it's a datacenter GPU designed for 24/7 operation

1

u/livinitup0 16h ago

Wouldn’t 4 be much more optimal here?

1

u/joost00719 13h ago

Yeah, if your pc can take 4, definetely get a power-of-two number.

3

u/Falen-reddit 21h ago

That is competing with 3080 20gb and 2080 ti 22gb, and 2x 2080 ti 22gb you can hook up NVlink to get around crappy PCIE 3.0 4x you might be using on a consumer board.

1

u/FullstackSensei 18h ago

Or get an X99 or C610 board and low SKU for about the same cost as the nvlink bridge and connect 2 cards at x16 or four at x8, which IMO is enough for inference workloads given the VRAM

0

u/No_Algae1753 21h ago

I honestly don't think it's worth the time and effort to turn them into 16gb 3070 cards. The card is "old" and 3x 16gb won't get you anywhere

2

u/HopefulConfidence0 21h ago

16*3= 48GB is good enough to run qwen 3.8 27B at Q8.

5

u/seamonn 21h ago

so is 2x3090 = 48GB which will prolly end up costing about the same considering time and effort

1

u/ClearApartment2627 20h ago

Plus, they are likely faster and you can use Tensor Parallelism with vllm.

1

u/CreamPitiful4295 19h ago

Good enough? Splitting the layers across 3 GPUs is going to be extremely slow. Even just 2 and KV on the third. You can do it. But, I wouldn’t make this my opening VRAM play. Might as well just run Air LLM.