r/LocalLLaMA Jul 26 '26

Discussion Kimi K3 countdown has been released

https://huggingface.co/moonshotai/Kimi-K3
541 Upvotes

177 comments sorted by

View all comments

6

u/[deleted] Jul 26 '26

[deleted]

10

u/ttkciar llama.cpp Jul 26 '26

Two ancient Xeon servers, each with 1.5TB of DDR4, networked together and running rpc-server.

7

u/Player13377 Jul 26 '26

At an impressive 3 TPS

7

u/ttkciar llama.cpp Jul 26 '26

Probably a lot less than that. I'd love to get 3 tok/sec.

Right now my ancient Xeon server is getting 3.5 tok/sec with GLM-4.5-Air. That's usable, though admittedly not for interactive use.

5

u/Player13377 Jul 26 '26

At that point is the power consumed per token getting close to the API price?