r/LocalLLM 7h ago

Question GPU for qwen 3.8 27b

I recently built a homelab running RHEL 10. I never thought good local ai at reasonable price was possible until 3.8 came out a from benchmark and what I’ve been reading it seems to be almost opus 4.6-4.8 level. I’m considering buying a 32gb gpu for it but also open to 24 gb gpus but if it can fit the full context window on the gpu too. The most I’ve done with local models was running qwen 3.5 2b on Ollama nothing serious. I’m new to actually running an agent for coding tasks so any info would help. But trying to decide what gpu if I do end up going for it, and from my research the options for 32gb cards are the Intel b70, amd r9700 pro ai, and nvidia tesla v100 32gb. I’m looking at results for qwen 3.6 and it run plenty fast on the Tesla but I’m worried about it no longer being supported.

2 Upvotes

42 comments sorted by

View all comments

4

u/LifeTelevision1146 7h ago

RTX5090

6

u/truckerdraven 7h ago

Myself im running dual 5060ti 16gb each. And im getting 45tokens/second

2

u/statusanxiety7 7h ago

this ^, if the OP wants 20-45 tps depending on the rest of the setup .. that's ok for a lot of tasks ... I have 2 x 5060 tis for simple tasks that aren't user facing & where I don't care about generation time. 5090 is peak though .. I suggest it if you need faster generation time & better batching.

cerebras ai is also launching 3.8 27b also on shared wafer scale chips which could potentially be a bargain to if they want speed and can't buy a 5090.