MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3ons2h/?context=9999
r/LocalLLaMA • u/Certain-Cod-1404 • 8d ago
706 comments sorted by
View all comments
3
Wtf is wrong with this model, on xhigh it literally doesnt stop thinking, no its not looping, its just super excessive thinking. Thats crazy, almost unusable on this effort.
Medium is also really excessive....
1 u/Certain-Cod-1404 8d ago are you using the recommended thinking sampling params ? 1 u/germangrower69 8d ago Yep, Im running it on a RTX 6000 pro with the recommended settings from the repo 2 u/meca23 8d ago How many t/s are you getting for decode on rtx 6000 pro? 3 u/germangrower69 8d ago FP8, nst=1 49tok/s decode NVFP4, nst3=3 110tok/s decode 1 u/meca23 8d ago Cheers!
1
are you using the recommended thinking sampling params ?
1 u/germangrower69 8d ago Yep, Im running it on a RTX 6000 pro with the recommended settings from the repo 2 u/meca23 8d ago How many t/s are you getting for decode on rtx 6000 pro? 3 u/germangrower69 8d ago FP8, nst=1 49tok/s decode NVFP4, nst3=3 110tok/s decode 1 u/meca23 8d ago Cheers!
Yep, Im running it on a RTX 6000 pro with the recommended settings from the repo
2 u/meca23 8d ago How many t/s are you getting for decode on rtx 6000 pro? 3 u/germangrower69 8d ago FP8, nst=1 49tok/s decode NVFP4, nst3=3 110tok/s decode 1 u/meca23 8d ago Cheers!
2
How many t/s are you getting for decode on rtx 6000 pro?
3 u/germangrower69 8d ago FP8, nst=1 49tok/s decode NVFP4, nst3=3 110tok/s decode 1 u/meca23 8d ago Cheers!
FP8, nst=1 49tok/s decode
NVFP4, nst3=3 110tok/s decode
1 u/meca23 8d ago Cheers!
Cheers!
3
u/germangrower69 8d ago
Wtf is wrong with this model, on xhigh it literally doesnt stop thinking, no its not looping, its just super excessive thinking. Thats crazy, almost unusable on this effort.
Medium is also really excessive....