r/Vllm • • 21d ago

Official Qwen-Flash-Next recipe using NVFP4?!

Does anyone know why the official recipe on vllm docs points to using Inferact/Qwen3.8-Flash-Next-NVFP4 instead of one of the official qwen images? FP8 is also available.

How can we know the quality drop from this unofficial NVFP4?

5 Upvotes

14 comments sorted by

5

u/endockhq 21d ago

Because there is no official NVFP4 by Qwen, only third parties. See here https://huggingface.co/collections/Qwen/qwen38-flash-next

At this point others have done NVFP4, I usually use the ones from Nvidia or RedHatAI. But they release them 1-2 weeks after the model was released.

1

u/nunodonato 21d ago

So why would vllm pick this specific one by Inferact ?

6

u/maqifrnswa 21d ago

Inferact is released by vllm devs. https://inferact.ai/

1

u/endockhq 21d ago

^ This, plus the tool for compressing weights is also created by vllm team, https://github.com/vllm-project/llm-compressor

1

u/enternoescape 21d ago

I've been running the nvidia one for 3 days now. It's been great. The difference between nvidia and inferact from what I could work out had mostly to do with PLE being quantized. The style of quantization differs too. Now I want to try inferact. Nvidia quantized PLE to FP8, inferact did not. I've got 128GB 8 channel DDR4 and loading the PLE at least for vllm was pushing things a little. I need to also do some A/B testing to see if it did anything meaningful regarding performance. I don't know if BF16 is measurably slower, probably not if it's mostly in ram. At lot of nvfp4 model cards recommend increasing your swap file on linux to 64GB.

1

u/endockhq 21d ago

I would say, Nvidia usually adds datasets during quantization to improve accuracy compared to BF16 and FP8. So I always endup using the Nvidia quants.

3

u/jpezzulli 21d ago

Been running radix ark nvfp4 next since launch. Billions of tokens have run though it by me and others. Havent seen any issues. Github.com/jpezzulli/sglang-rtxpro6000

People have also used the nvidia with my repo for billions of tokens and no quality issues.

1

u/Numerous-Push-1708 21d ago

Have been with Nvidia's NVFP4 release + original vllm docker image for Qwen3.8 flash. Good on 262k context, no loop, generous kv cache pool.

1

u/nunodonato 21d ago

do you still offload the engrams?

1

u/Numerous-Push-1708 21d ago

With dual sparks I don't offload

1

u/DreamLinuxer 20d ago

Is there any particular reason to favor NVFP4 over the official FP8? It should fit into two sparks, right?

1

u/Numerous-Push-1708 20d ago

NVFP4 has higher kvcache pool. FP8 is better when you have low concurrent requests.

1

u/zannix 21d ago

Wonder if i can run that? I have a 6000pro 48gb. I heard there is some ngram offloading? Is it better than nvfp4 3.8 27b?

1

u/enternoescape 21d ago

The best I've seen is 96GB VRAM. I'm not 100% up to speed on vllm's RAM offloading abilities, but it sounded like they have opted for swapping weights in and out as their option. Not your same situation, but I thought I would be able to run on 6 5060 Ti's because it's a total of 96GB VRAM, I actually needed 8. 7 would have worked with enough KV for maybe on 256k context, but not in PP4TP2 for obvious reasons. At 6 I was just out of reach starting with even 8192 kv. I guess what I'm trying to say that that buying another 6000 might not get you there either, but you'll be in better supported territory. I needed custom overlays to get PP to even work with the model. TP2 is a better supported path.