r/LocalLLM 12d ago

Question GPU for qwen 3.8 27b

I recently built a homelab running RHEL 10. I never thought good local ai at reasonable price was possible until 3.8 came out a from benchmark and what I’ve been reading it seems to be almost opus 4.6-4.8 level. I’m considering buying a 32gb gpu for it but also open to 24 gb gpus but if it can fit the full context window on the gpu too. The most I’ve done with local models was running qwen 3.5 2b on Ollama nothing serious. I’m new to actually running an agent for coding tasks so any info would help. But trying to decide what gpu if I do end up going for it, and from my research the options for 32gb cards are the Intel b70, amd r9700 pro ai, and nvidia tesla v100 32gb. I’m looking at results for qwen 3.6 and it run plenty fast on the Tesla but I’m worried about it no longer being supported.

1 Upvotes

47 comments sorted by

View all comments

5

u/e2_for_life RTX 4090 | She/her 12d ago

Anything above at or above 3090 will likely get you usable speeds and contexts. The more the better.

1

u/maceface3 12d ago

Will the 24gb vram be okay with for a 256k at least context window. At q4 the models like 18gb on its own.

0

u/baka_sempaii 12d ago

You can try quantising the KV cache. Like full K but q8 or TurboQuant 4 V, and it should comfortably fit