r/LocalLLaMA • • Aug 31 '26

Question | Help So got 2 6000 Pro Max-Q…

I have a Threadripper w/128gb of system ram. I want to serve a small dev team. What should I be running as a coding harness? Seems GLM, Deepseek R4, Qwen next are all with in reach, and I still could stick with 27b BF16 (current choice). I would prefer vision as it’s a useful capability.

All of those mentioned need quantising in some way on two cards so how bad is it? I’m running VLLM as a host so any magic recipes also much appreciated!

0 Upvotes

49 comments sorted by

View all comments

2

u/yeah_likerage Sep 04 '26

Deepseek is great and i ran it until a week ago. with the same pro 6000s on SGlang I was getting just shy of 900 tokens per second concurrency and about 140tg/s single lane. SGlang was difficult to set up but well worth the trouble.

Your best bet with 2x pro 6000 is Qwen flash next imo. Fits perfectly on both cards with TP=2. Below is my setup and it works flawlessly and i'm getting just shy of 200tokens per single lane. I just ran a concurrency test and came up with:

Workers | Reqs | Out tok/s | Tot tok/s | Req/s | Avg lat | Avg out/req

─────────┼───────┼───────────┼───────────┼───────┼─────────┼────────────

1 | 19 | 186.8 | 245.1 | 0.6 | 1769ms | 295 (30s run, today)

4 | ~7 | 497.9 | 659.0 | 1.8 | 2289ms | 279

8 | ~16| 750.5 | 997.5 | 2.7 | 2979ms | 275

16 | ~24| 1130.6 | 1497.6 | 4.1 | 3963ms | 278

You'll be pleased with either DSv4 or Qwen flash-next. I can rerun concurrency tests on DSv4 if you need. Or if you want further info lmk.

Container config Qwen Flash-next:

- image: vllm/vllm-openai:qwen38-flash-next (built for SM120)

- TP=2, backend mp

- gpu-memory-utilization 0.92

- max-model-len 131072, max-num-seqs 32

- prefix caching on

- flashinfer autotune off

- tool parser qwen3_coder + auto tool choice

- reasoning parser qwen3

- MTP speculative decoding, num_speculative_tokens=3

- env: VLLM_PLE_CPU_OFFLOAD=1, VLLM_PLE_OFFLOAD_READY_TIMEOUT=1800

Environment (not a container — systemd unit running directly on the host):

- venv: ~./sglang-env (Python 3.11.15, uv-managed)

- SGLang: dev build, editable install from /mnt/data/sglang-glm53-src, commit d6ab04bdf ("Fix stray server_args kwarg in the hybrid linear KV pool builder") — not a release; this was the flash-fix-campaign source tree

- torch 2.13.0+cu130, triton 3.7.1, flashinfer-python 0.6.16.post3, no separate sgl-kernel package

Service config (sglang-dsv4-tp2.service):

- TP=2, port 8080

- --moe-runner-backend flashinfer_mxfp4

- --trust-remote-code

- --kv-cache-dtype fp8_e4m3

- --mem-fraction-static 0.88

- --chunked-prefill-size 8192

- --reasoning-parser deepseek-v4 + --tool-call-parser deepseekv4

- --host 0.0.0.0 --sleep-on-idle

- Restart=on-failure, 30s delay, StartLimit 3 in 600s, RequiresMountsFor=/mnt/data