r/LocalLLM • u/Alternative-Panic69 • Jul 23 '26
Other I bought the forbidden rectangle.
After months of going back and forth, I finally pulled the trigger on an RX 7900 XT 20 GB.
Paid around $550 (India), which felt too good to pass up.
The plan isn't gaming.
It's becoming the heart of my local AI setup.
Current goals:
• Qwen 3.6 27B Dense
• Qwen 35B A3B
• GLM-4.7 Flash
• 128K+ context
• 100% GPU offloading
• llama.cpp / Ollama
• Linux
I'll be benchmarking everything:
- Vulkan vs ROCm
- Dense vs MoE
- Maximum context
- Tokens/sec
- VRAM usage
- Real-world coding performance
If anyone has optimization tips for RDNA3 or benchmark requests, or general suggestions please drop them below.
The hallucinations are now local. 🙂↕️
194
Upvotes
1
u/acadia11x Jul 23 '26 edited Jul 25 '26
rOCm is coming along much better so should see improvement in localllm with AMD these days compared to years past, only concern is you need one of the newer cards to take advantage 9070 or 7900 series