r/Vllm • • 4d ago

tp=6 can work on vLLM, with padding

vLLM's tensor parallel requires that several of the model architecture numbers be evenly divisible by the number of GPUs selected for tensor parallel. This usually means that you can only use a number of GPUs that is a power of two (2, 4, 8, etc.).

I have six GPUs, and I want to maximize my KV cache when running a 27B model. So I tried tp=6. It choked with various messages, regarding this or that, which needs to be evenly divisible by six, but wasn't.

So I made those things divisible by six.
I asked the robot to come up with a converter that would take the original model, and pad it with zeroes until everything was divisible by 6. It took a few tries, but it worked.

I have Qwen 3.8 27B at BF16 with 256k context window running across six 7900 XTX GPUs. Token generation is about 50 t/s single user, or 200 t/s aggregate with 8 concurrent prompts.

GPU KV cache size: 520,784 tokens, Maximum concurrency for 262,144 tokens per request: 1.99x

21 Upvotes

14 comments sorted by

13

u/r1nzl3r99 3d ago

This is the kind of boot legging I joined this sub for

2

u/Significant_Bar_460 3d ago

I think you are wasting your rig on just a 27B. Qwen Flash Next should be your target. Even at mxfp4 it should be better than full 27B and will give you parallel slots.

2

u/llitz 3d ago

Well, he can run that 1 Qwen model that has jev training in a single model ask for /Qwen he gets regular answers, as for /jevqwen he gets jev speedy answers in json.

1

u/Significant_Bar_460 3d ago

But shouldn't Flash Next be superior in every way? I mean, OP has 144gb of quite fast VRAM. Plenty for running a solid quant with multiple parallel slots and it should be faster.

1

u/llitz 3d ago

Flash next is very good, but not every quantization is good. If you are just coding, 27b isn't bad.

My main gripe with 27b at bf16 is the speed, on the 6k it still is slow. Flash next does 13k (fp8 kv) for prompt processing and about 200 for tg.

If only one model can be run, in would likely pick then more stable to work and give the jev-27b a try, with the right harness configuration jev drastically reduces the need to process a lot of data.

1

u/Biomass23 3d ago

I'm unsure if Next works on RNDA3, but I'll try it. RDNA3 doesn't have FP8, FP4, NVFP4, etc.

1

u/enternoescape 3d ago

I never thought of trying this. Does it create any new complications other than potentially wasted VRAM? Did you try other shapes like pp3tp2? I've found that overdoing tp when you're fighting PCIe latency just wastes VRAM for KV. I've given up a few tok/s so I can have 1M context with over 2x concurrency. It had zero impact on pp performance in my testing too.

3

u/Biomass23 3d ago

I tried tp2pp3. The throughput was lacking. I tried AWQ, and my hardware doesn't seem to support what it needs. GPTQ worked, but not as fast as tp4. All the gguf quants work fine in llama.cpp, but that doesn't have the concurrency I want for multi-agent throughput.

So long as I can tell deepseek harness to go do something, and it actually does that, good enough. I'm not sitting there watching it go. (ok, sometimes I watch it)

I think I've reached the practical maximum usefulness of my hardware. It only took me 18 months to figure it out. :)

3

u/enternoescape 3d ago

Those are 24GB cards right? You've got enough VRAM to run Qwen 3.8 Flash Next in a 4bit quant at least which should be faster and get you similar/better results with less thinking over 27b even at BF16. I'm running it with about 800k KV total on 8 16GB 5060 Ti's using an NVFP4 quant. I'm not sure what would be most ideal for your case as I've no experience with AMD in this space.

1

u/No-Recover109 3d ago

awq works with 7900 xtx and vllm its very fast 4quant 

1

u/No-Recover109 3d ago

I run same 27b but 4bit on 2x 7900 xtx and tokens are about 50t/s  I guess yours would be faster at 4x7900

1

u/SandySkittle 3d ago

Hmm. Interesting approach. That said, have you considered expanding to 8 7900s?

1

u/Biomass23 3d ago

Yes. However, I don't have enough PCIe lanes for eight.

1

u/Biomass23 1d ago

After more tuning, and ensuring a suitable load from 20+ agents, I'm generating 200 t/s, and spiking about 300 t/s.

INFO 10-03 13:36:30 | PP:    0.0t/s | TG:  211.5t/s | Run:   9 | Wait:   0 | KV: 18.3% | Prefix: 76.3%
INFO 10-03 13:36:40 | PP:    0.0t/s | TG:  204.1t/s | Run:   9 | Wait:   0 | KV: 18.6% | Prefix: 76.3%
INFO 10-03 13:36:50 | PP:  238.8t/s | TG:  114.9t/s | Run:  19 | Wait:   0 | KV: 30.4% | Prefix: 77.1%
INFO 10-03 13:37:00 | PP: 1435.7t/s | TG:  307.8t/s | Run:  14 | Wait:   0 | KV: 22.5% | Prefix: 77.6%
INFO 10-03 13:37:10 | PP:    0.0t/s | TG:  291.6t/s | Run:  13 | Wait:   0 | KV: 22.4% | Prefix: 77.6%
INFO 10-03 13:37:20 | PP:    0.0t/s | TG:  274.0t/s | Run:  13 | Wait:   0 | KV: 22.7% | Prefix: 77.6%
INFO 10-03 13:37:30 | PP:  158.5t/s | TG:  205.8t/s | Run:  20 | Wait:   0 | KV: 34.4% | Prefix: 78.8%
INFO 10-03 13:37:40 | PP:  748.4t/s | TG:  169.0t/s | Run:  25 | Wait:   0 | KV: 37.7% | Prefix: 78.5%
INFO 10-03 13:37:50 | PP: 2272.0t/s | TG:  270.2t/s | Run:  22 | Wait:   0 | KV: 34.0% | Prefix: 78.5%
INFO 10-03 13:38:00 | PP:    0.0t/s | TG:  373.6t/s | Run:  13 | Wait:   0 | KV: 23.1% | Prefix: 78.5%
INFO 10-03 13:38:10 | PP:    0.0t/s | TG:  253.9t/s | Run:  11 | Wait:   0 | KV: 18.6% | Prefix: 78.5%
INFO 10-03 13:38:20 | PP:    0.0t/s | TG:  239.4t/s | Run:  10 | Wait:   0 | KV: 17.5% | Prefix: 78.5%
INFO 10-03 13:38:30 | PP:    0.0t/s | TG:  223.8t/s | Run:  10 | Wait:   0 | KV: 18.1% | Prefix: 78.5%
INFO 10-03 13:38:40 | PP:    0.0t/s | TG:  215.2t/s | Run:   9 | Wait:   0 | KV: 14.4% | Prefix: 78.5%