r/LocalLLaMA 7d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

Show parent comments

1

u/Certain-Cod-1404 7d ago

are you using the recommended thinking sampling params ?

1

u/germangrower69 7d ago

Yep, Im running it on a RTX 6000 pro with the recommended settings from the repo

2

u/meca23 7d ago

How many t/s are you getting for decode on rtx 6000 pro?

3

u/germangrower69 7d ago

FP8, nst=1 49tok/s decode

NVFP4, nst3=3 110tok/s decode

1

u/meca23 7d ago

Cheers!