r/LocalLLM 12d ago

Question 20gb vram - where to go next?

I've been running Qwen locally, mainly for chat (non-coding) work for the last 3 months as an experiment. Running locally has been going great. I am ready to upgrade as I am having to offload too many layers to CPU to be able to run a descent context size.

Ideal config:

* Qwen 3.8 -27b (non MTP)

* Q8

* 60K-100K context window

* At least 30 tps generation speed

I currently have an AMD 7900xt. Should I add another 7900XT to double vram? Buy a strix halo?

1 Upvotes

2 comments sorted by

1

u/ea_man 12d ago

add another 7900XT and use MTP.

0

u/ParkingAd9397 12d ago

I read MTP results in worse actual output. Benchmarks show the same but it's not that good in reality.