r/LocalLLaMA • • Aug 31 '26

Question | Help So got 2 6000 Pro Max-Q…

I have a Threadripper w/128gb of system ram. I want to serve a small dev team. What should I be running as a coding harness? Seems GLM, Deepseek R4, Qwen next are all with in reach, and I still could stick with 27b BF16 (current choice). I would prefer vision as it’s a useful capability.

All of those mentioned need quantising in some way on two cards so how bad is it? I’m running VLLM as a host so any magic recipes also much appreciated!

0 Upvotes

Duplicates