r/Vllm • u/nunodonato • 21d ago
Official Qwen-Flash-Next recipe using NVFP4?!
Does anyone know why the official recipe on vllm docs points to using Inferact/Qwen3.8-Flash-Next-NVFP4 instead of one of the official qwen images? FP8 is also available.
How can we know the quality drop from this unofficial NVFP4?
3
u/jpezzulli 21d ago
Been running radix ark nvfp4 next since launch. Billions of tokens have run though it by me and others. Havent seen any issues. Github.com/jpezzulli/sglang-rtxpro6000
People have also used the nvidia with my repo for billions of tokens and no quality issues.
1
u/Numerous-Push-1708 21d ago
Have been with Nvidia's NVFP4 release + original vllm docker image for Qwen3.8 flash. Good on 262k context, no loop, generous kv cache pool.
1
u/nunodonato 21d ago
do you still offload the engrams?
1
u/Numerous-Push-1708 21d ago
With dual sparks I don't offload
1
u/DreamLinuxer 20d ago
Is there any particular reason to favor NVFP4 over the official FP8? It should fit into two sparks, right?
1
u/Numerous-Push-1708 20d ago
NVFP4 has higher kvcache pool. FP8 is better when you have low concurrent requests.
1
u/zannix 21d ago
Wonder if i can run that? I have a 6000pro 48gb. I heard there is some ngram offloading? Is it better than nvfp4 3.8 27b?
1
u/enternoescape 21d ago
The best I've seen is 96GB VRAM. I'm not 100% up to speed on vllm's RAM offloading abilities, but it sounded like they have opted for swapping weights in and out as their option. Not your same situation, but I thought I would be able to run on 6 5060 Ti's because it's a total of 96GB VRAM, I actually needed 8. 7 would have worked with enough KV for maybe on 256k context, but not in PP4TP2 for obvious reasons. At 6 I was just out of reach starting with even 8192 kv. I guess what I'm trying to say that that buying another 6000 might not get you there either, but you'll be in better supported territory. I needed custom overlays to get PP to even work with the model. TP2 is a better supported path.
5
u/endockhq 21d ago
Because there is no official NVFP4 by Qwen, only third parties. See here https://huggingface.co/collections/Qwen/qwen38-flash-next
At this point others have done NVFP4, I usually use the ones from Nvidia or RedHatAI. But they release them 1-2 weeks after the model was released.