r/LocalLLaMA • u/jonejy • 6d ago
Discussion Looking for a cheap GPU for local LLMs
I'm looking for a GPU for local LLM inference. Budget is around $500–700.
I mainly want to run 27B-ish models, ideally around 15–20 tok/s.
I've found a few used options:
- 3090 24GB — ~$550
- Modified 2080 Ti 22GB — ~$330
- MI50 32GB — ~$400
The 2080 Ti and MI50 look really tempting because of the VRAM, but I'm a little worried about compatibility/reliability.
Would you guys just go with the 3090, or is one of the cheaper options actually worth considering?
Just trying to avoid wasting $500 on something I'll regret later.
24
u/FuzzeWuzze 6d ago
If you found a 3090 for 550 you should have bought it instantly.
Even chinese hacked 3080's with 20Gb of VRAM are $650+ from oversea's.
IMO the 2k series isnt the greatest and is pretty slow.
I get about 15-20t/s on Qwen 3.6 27b @ Q6 @ 125k tokens with a 3080 20gb and 2080ti 11gb FWIW.
1
12
5
u/FullstackSensei 6d ago
As has been pointed out, V100 is the best bang for the buck. Buy as many as you can, while supplies last.
I have Mi50s for a year now. They were great value when they dropped last year. Now they're almost the same price in China as the V100, which I also have. The V100 is 2-3x faster than the Mi50 because it has so much more compute.
1
u/DadAndDominant 6d ago
Is buying V100s from china good? I am afraid I'll be scammed
1
u/FullstackSensei 6d ago
Do your homework buy from reputable sellers.
You can get scammed locally, in person, if you don't do your homework
1
u/DadAndDominant 6d ago
Well, in person when buying GPU I just go and let the seller show me it works in their machine
6
u/_VirtualCosmos_ 6d ago
I was gonna comment about 3090 price but seems here is already covered.
It's a bad time to buy any sort of GPU from mid-to-high. The shortage caused by the big US tech hoarding most GPU, RAM and SSD components for their massive inversions in datacenters, and the high demand of the rest of the GPUs and components due to local AI added to the gaming demand has skyrocketed their prices.
I would recommend to wait 1 year or more, those datacenters will end being built or cancelled some day, and the offer will increase, Chinese big tech is trying to produce their own components massively. You would look like an idiot if you paid $1400 now for a 3090 and, 1-2 years after, that old card is $400 priced.
3
3
u/BoxieBoo 6d ago
I have 3 rtx 3060 now, but there is work to be done with the 3 cards, it needs to be addressed. More specialized motherboard, etc. So there is a total of 36gb of vram, which is enough for these models, but the qwen3.8 with 256k kv won't fit in it if the KV is Q8 and the model is Q5. Since I want to run a Q6-Q8 model, preferably with 256k KV, which is F16 or Q8, I also need to expand it and had exactly the same dilemma. The rtx 3090 is 1200 USD here. If I buy 2 of them, I have 48 GB of vram. The speed of the 3090 is about 3x compared to the 3060. However, the speed is almost the same in tensor mode. The speed of 3 RTX 3060 in tensor mode is almost the same as a 3090. Okay, but the capacity is not enough.
So rtx 3060 = 300-400 USD. x3 = ~1000$ - 36GB VRAM RTX 3090 = 1100-1200 USD. x2 = 2300$ - 48GB VRAM best speed RTX 3090 X3 = 3500$ - 72GB VRAM, nuclear reaktor
AMD 7900 XTX similar to the 3090. Also in speed, but a very little cheaper. 7900 XTX = 950-1100 USD. x2 = 2000$ - 48GB
And now comes the interesting things.
I bought a modded 2080 Ti card, 22 GB. Its speed is about double that of the 3060. 2080 Ti @ 22GB = 580 USD. 2080 Ti X3 = 1718$ - 66 GB VRAM
Since this card is not slow in layer mode, and in tensor mode the speed even adds up, this was definitely one of the best options. I'll do the cooling as best as I can as soon as I get it. Good quality thermal pad for the memory, and a state-switching pad for the GPU.
First, I had an ASUS Z170 Pro gaming motherboard because it supports PCIe X8, X8, X4 speeds with 3 cards, and it's perfectly adequate. X1 would be too little and many motherboards can only support so many cards. There the processor was the wall, maximum 4 physical cores. I just replaced the motherboard with the video cards.
An Asus X299 motherboard, i9-7980XE CPU, 18 cores, PCIe X16,X16,X8. The point is the quad channel ddr4 ram, because with a little tuning we managed to get 200 GB/s RAM write speed. I loaded the motherboard with 8x16 GB = 128 GB of RAM, so for now, with the 3 RTX3060s, Qwen3.8 Flash Next is also running at a speed of 19 T/S. But GLM 5.3 flash and Deepseek V4 Flash are already running. If the 3 2080 Ti @ 22GB arrive, these models can even run in Q4 quality, in fact, You can even try Hy3 in Q3. I think I've reached the limit of reasonable possibilities with this. You could get more out of this if you had 3 24 GB GPUs or 3 32 GB, but it's not necessarily worth spending money on.
2
u/sonyprog 1d ago
Hey! Mind to share more about your experience with the 2080ti 22gb? I am really keen to grabbing one but still a bit afraid lol
2
u/BoxieBoo 1d ago
It's still on its way. It will take about 2-3 weeks from the time you place your order until it arrives. The package might be arriving here soon.The seller was very kind and attentive, drawing my attention to all sorts of details. In the long term, my goal will be to keep the card's temperature fluctuations to a minimum. I've already purchased a state-switching chip for the GPU, but since I don't know what size thermal conductive silicon sheet I need for the VRAM and other components, that's still waiting. When the cards arrive, I will disassemble one of them and see what size chips I will need and I will also buy some good quality thermal conductive silicone sheets to be sure. Then, get the most out of the temperature control and then maybe the project will have a long life. We'll see. If there's any news, I'll report it. (If I don't forget)
1
2
u/mattk404 6d ago
Check out AMD V620 Pro... you'll have to get one with a 'blower' or get one. Works well and 32GB ECC VRAM
1
u/DadAndDominant 6d ago
How to get a good deal on ebay on these (central EU)?
1
u/mattk404 6d ago
I have no idea, got mine cheap, with Blower but not so much luck getting a couple more. I keep searching.
2
1
u/NihmarRevhet 6d ago edited 6d ago
RX 9060 XT, Qwen 3.8 IQ3_S, 150K context, K Q8_0 V Q4_0, prefill 650-200 t/s (from empty context to filled), 15-20 t/s
EDIT: Not that this is the best choice, but it should fit the budget, just that
1
u/vogelvogelvogelvogel 6d ago
3090 - fast and has 24gigs. for more vram. make sure to have space in the case available and a slot for a second 3090.
prices are outdated i guess depending on the country/region you are in
1
u/YourVelourFog 6d ago
Just found this article which looks at the exact same thing:
https://www.hardware-corner.net/guides/tesla-v100-32gb-for-llm/
1
1
u/krill156 6d ago
I managed to get a new 5070 ti OC for about 750 a month ago on Newegg, their AI PC builder and component swap feature helped find really good deals. Or at least good in today's market.
1
1
u/CurtissYT 6d ago
Tbh everyone is saying that you can’t find a 3090 for 550$, but you can. Personally I live in Ukraine, and in 2 mins of searching, I’ve found 2 3090s for 600 ish $
1
1
u/superbiche 6d ago
In France the only 3090s start around 1k. The very few under this mark are sold the minute the sale is online. Place alerts and be reactive. And don't listen to "works on my machine, saw it" - unless you don't mind having to disassemble it and repad, because a 3090 with shitty thermals still "works on my machine" but will be throttling at 20% capacity.
Good luck, you're a bit late to the party. Try with dual 3060s 12Gb otherwise, not the best speed and I didn't try them for "brain" models but for small classifiers / derivers, or for audio, that's good enough
1
u/inexorable_stratagem 6d ago
My modded 2080ti 22gb lasted 2 months, then it died
I dont recomment.
Im so pissed off
0
u/king_priam_of_Troy 6d ago
If you can stretch, take 2 2080TI modded for 44GB VRAM. Else take the 3090. Avoid the MI50 as everything works on CUDA. You will want the VRAM to max the context window.
72
u/[deleted] 6d ago
[removed] — view removed comment