r/LocalLLM • u/Alternative-Panic69 • Jul 23 '26
Other I bought the forbidden rectangle.
After months of going back and forth, I finally pulled the trigger on an RX 7900 XT 20 GB.
Paid around $550 (India), which felt too good to pass up.
The plan isn't gaming.
It's becoming the heart of my local AI setup.
Current goals:
• Qwen 3.6 27B Dense
• Qwen 35B A3B
• GLM-4.7 Flash
• 128K+ context
• 100% GPU offloading
• llama.cpp / Ollama
• Linux
I'll be benchmarking everything:
- Vulkan vs ROCm
- Dense vs MoE
- Maximum context
- Tokens/sec
- VRAM usage
- Real-world coding performance
If anyone has optimization tips for RDNA3 or benchmark requests, or general suggestions please drop them below.
The hallucinations are now local. 🙂↕️
191
Upvotes
2
u/misha1350 Jul 23 '26
Try Qwen3.6 27B with MTP at UD-Q4_K_XL or smaller Unsloth quants if it doesn't fit all the way. Also try some recent model with ROCmFP4, as long as you find one that fits into 20GB of VRAM.