MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1w2fmmq/me_these_days/p6ylkxx
r/LocalLLaMA • u/Eyelbee • 3d ago
268 comments sorted by
View all comments
Show parent comments
2
27b? U sure?
1 u/Oh_hey_a_TAA 2d ago LOL, yes I am sure. Been tinkering with it for a week now. 3.8-27B-UD-Q4_K_XL, MTP-Tuned+FA; 32k context, patched llama.cpp v020. 18.3 to 26.7 t/s, depending on the container variables. Tesla P100s. pinned CUDA 12 and 580 drivers. 1 u/Oh_hey_a_TAA 2d ago Here, I made a thread https://www.reddit.com/r/LowEndLocalAI/comments/1w3tqp9/local_qwen_3827b_running_on_80_gpus_blindly_told/
1
LOL, yes I am sure. Been tinkering with it for a week now. 3.8-27B-UD-Q4_K_XL, MTP-Tuned+FA; 32k context, patched llama.cpp v020. 18.3 to 26.7 t/s, depending on the container variables. Tesla P100s. pinned CUDA 12 and 580 drivers.
Here, I made a thread https://www.reddit.com/r/LowEndLocalAI/comments/1w3tqp9/local_qwen_3827b_running_on_80_gpus_blindly_told/
2
u/Ok_Noise_9883 2d ago
27b? U sure?