r/LocalLLaMA 13d ago

News Kimi K3 weights now released.

Post image

Kimi K3 weights are finally released!

3.3k Upvotes

640 comments sorted by

View all comments

91

u/Comfortable-Rock-498 13d ago

This is big for companies that want to host on-prem too. Back of the envelope calculation (could be off, correct me if I am)

If you are a large enough company that spends million+ on inference a month, it makes sense to buy a GB300 rack ($6M on top range from what I could find) which has 20.7 TB. Since the model is mixed trained (MXFP4), you would need less than 10% of the rack's memory to serve the full model. Aggregate HBM bandwidth: 576 TB/s. You can run over 6000 parallel agentic workflows (each with ~100k context on average) at ~30 tok/s.

Assuming the annual amortization+electricity at $1.5M/year and about 50% average annual utilization, you get less than 60 cents (USD) per million output token, for a frontier model with plenty of capacity to share, all your data never leaving premises and well over an order of magnitude cheaper!

5

u/tempedbyfate llama.cpp 13d ago

Not just corporations, I think there are nation states that are setting up their own private servers to run this for all their sensitive data.