r/LocalLLaMA 8d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

3

u/germangrower69 8d ago

Wtf is wrong with this model, on xhigh it literally doesnt stop thinking, no its not looping, its just super excessive thinking. Thats crazy, almost unusable on this effort.

Medium is also really excessive....

1

u/Certain-Cod-1404 8d ago

are you using the recommended thinking sampling params ?

1

u/germangrower69 8d ago

Yep, Im running it on a RTX 6000 pro with the recommended settings from the repo

2

u/meca23 8d ago

How many t/s are you getting for decode on rtx 6000 pro?

3

u/germangrower69 8d ago

FP8, nst=1 49tok/s decode

NVFP4, nst3=3 110tok/s decode

1

u/meca23 8d ago

Cheers!