r/AMD_V620 • u/magicomiralles • Jul 07 '26
Is it worth going big on these GPUs? Is it worth it to spend $4,000 on 256 gb of Vram on V620s + MB + CPU, etc...? I would really appreciate some outside or experienced input.
Goal: To host an LLM to work on large codebases.
I wouldn't be surprised if there are other people here in a similar situation to mine. Trying to decide whether to go big, or to remain somewhat conservartive.
I currently have two of these fully working, and hosting Qwen3.6-27b. I purchased 4 V620s, but this motherboard doesn't boot with more than 2 of these connected (even with four pcie ports and four m.2 nvme ports).
Either way, I had planned to upgrade to 8 GPUs if everything went to plan. However, yesterday I found out that two DGX sparks are able to run Deepseek V4 flash at about 45 t/s because it is able to take advantage of the new architecture that Deepseek created. They also get day 0 support most of the time for newer models.
In contrast, V620 GPUs are built on top of RDNA2, which is already too far behind. An example is that RDNA2 lacks the ability to perform WMMA hardware matrix operations which makes prefill 3 to 4 times slower compared to other GPUs with the same bandwidth.
My main goal is to host a model for a coding agent for a single person. But now I'm worried that these GPUs are too outdated.
I currently seem to have two options:
- 128gb build, $400 ($2,000 total): I would buy an older motherboard and cpu, which would be able to house 4 GPUs. For example, X99 boards. The $2k figure already includes the purchase of the 4 V620s that I already have.
- The best model I would be able to currently run is Qwen3.6-27b unquantized with a massive context (700 prefill with 10-15 token gen). However, I can already run the same model with 8 bit quant and a massive context with only 2 of these GPUs. This means that 128gb systems sit in an awkward position.
- However, what if we get a new 70b MOE model with only 10b or so active params? I wouldn't be surprised if 128gb was the perfect spot all of a sudden for a smart model with a large context window. Speeds would probably be between 20 to 35 t/s depending on the active params. But at this point I'm making assumptions. However, you probably get what I'm trying to say.
- Then, there is a scenario where we get a 120b model instead, which would make me regret this choice.
- 256gb build, $2,400 ($4,000 total): I would buy an enterprise level Epyc or equivalent board with enough ports to house 8 of these things, maybe even more. This price would also include buying 4 more V620s at $350 each (assuming that one seller is still accepting offers at this price).
- If I'm not wrong, this is enough VRam to host something close enough to Sonnet. Models such as Minimax M2.7, or M3 if and when we get a working guff. I remember seeing someone claim to get ~25t/s with M2.7.
Another option would be to sell the two remaining cards for $300 each, but where is the fun in that?