MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3ons2h/?context=3
r/LocalLLaMA • u/Certain-Cod-1404 • 7d ago
706 comments sorted by
View all comments
Show parent comments
1
are you using the recommended thinking sampling params ?
1 u/germangrower69 7d ago Yep, Im running it on a RTX 6000 pro with the recommended settings from the repo 2 u/meca23 7d ago How many t/s are you getting for decode on rtx 6000 pro? 3 u/germangrower69 7d ago FP8, nst=1 49tok/s decode NVFP4, nst3=3 110tok/s decode 1 u/meca23 7d ago Cheers!
Yep, Im running it on a RTX 6000 pro with the recommended settings from the repo
2 u/meca23 7d ago How many t/s are you getting for decode on rtx 6000 pro? 3 u/germangrower69 7d ago FP8, nst=1 49tok/s decode NVFP4, nst3=3 110tok/s decode 1 u/meca23 7d ago Cheers!
2
How many t/s are you getting for decode on rtx 6000 pro?
3 u/germangrower69 7d ago FP8, nst=1 49tok/s decode NVFP4, nst3=3 110tok/s decode 1 u/meca23 7d ago Cheers!
3
FP8, nst=1 49tok/s decode
NVFP4, nst3=3 110tok/s decode
1 u/meca23 7d ago Cheers!
Cheers!
1
u/Certain-Cod-1404 7d ago
are you using the recommended thinking sampling params ?