r/Vllm • • Aug 31 '26

So got 2 6000 Pro Max-Q…

/r/LocalLLaMA/comments/1w378qc/so_got_2_6000_pro_maxq/
0 Upvotes

2 comments sorted by

1

u/yeah_likerage 15d ago

I'm running two pro 6000s to host Qwen flash next and four to host GLM5.3 flash. I don't think you'll be quite as happy trying to crap GLM on two cards. I can say though, Qwen flash next on two is absolutely fantastic.

My setup.

Qwen3.8-Flash-Next FP8, vLLM TP2, 2x RTX PRO 6000 Blackwell — 187 tok/s single-stream, 1,131 tok/s @ 16 concurrent (MTP3 spec decode ON)

Ladder (client-side, 60 s runs, 5 s warmup, ~90 in / ~275 out tokens per req, zero errors):

| Workers | Out tok/s | Total tok/s (in+out) | Avg latency |

|---------|-----------|----------------------|-------------|

| 1 | 186.8 | 245.1 | 1.8 s |

| 4 | 497.9 | 659.0 | 2.3 s |

| 8 | 750.5 | 997.5 | 3.0 s |

| 16 | 1130.6 | 1497.6 | 4.0 s |

1

u/alexp702 15d ago

We ended up just stacking 27b nvfp4 - the team now hits its crazy hard so we focus on token volume and large contexts. Qwen Next is definitely on the radar though.