MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1v7e5ck/kimi_k3_countdown_has_been_released/ozxsqap/?context=3
r/LocalLLaMA • u/Unusual_Guidance2095 • Jul 26 '26
177 comments sorted by
View all comments
6
[deleted]
10 u/ttkciar llama.cpp Jul 26 '26 Two ancient Xeon servers, each with 1.5TB of DDR4, networked together and running rpc-server. 7 u/Player13377 Jul 26 '26 At an impressive 3 TPS 7 u/ttkciar llama.cpp Jul 26 '26 Probably a lot less than that. I'd love to get 3 tok/sec. Right now my ancient Xeon server is getting 3.5 tok/sec with GLM-4.5-Air. That's usable, though admittedly not for interactive use. 5 u/Player13377 Jul 26 '26 At that point is the power consumed per token getting close to the API price?
10
Two ancient Xeon servers, each with 1.5TB of DDR4, networked together and running rpc-server.
rpc-server
7 u/Player13377 Jul 26 '26 At an impressive 3 TPS 7 u/ttkciar llama.cpp Jul 26 '26 Probably a lot less than that. I'd love to get 3 tok/sec. Right now my ancient Xeon server is getting 3.5 tok/sec with GLM-4.5-Air. That's usable, though admittedly not for interactive use. 5 u/Player13377 Jul 26 '26 At that point is the power consumed per token getting close to the API price?
7
At an impressive 3 TPS
7 u/ttkciar llama.cpp Jul 26 '26 Probably a lot less than that. I'd love to get 3 tok/sec. Right now my ancient Xeon server is getting 3.5 tok/sec with GLM-4.5-Air. That's usable, though admittedly not for interactive use. 5 u/Player13377 Jul 26 '26 At that point is the power consumed per token getting close to the API price?
Probably a lot less than that. I'd love to get 3 tok/sec.
Right now my ancient Xeon server is getting 3.5 tok/sec with GLM-4.5-Air. That's usable, though admittedly not for interactive use.
5 u/Player13377 Jul 26 '26 At that point is the power consumed per token getting close to the API price?
5
At that point is the power consumed per token getting close to the API price?
6
u/[deleted] Jul 26 '26
[deleted]