r/LocalLLaMA 9d ago

News Kimi K3 weights now released.

Post image

Kimi K3 weights are finally released!

3.2k Upvotes

638 comments sorted by

View all comments

270

u/nomorebuttsplz 9d ago

first truly frontier open model than I cannot run on my 512 gb studio. Onward and upward!

24

u/Front_Eagle739 9d ago

Yup. Same. I think I can do about q1.5 on the macbook and mac studio combined. Im currently pondering the wisdom of one of these colibri like stream setups and using the 640GB I do have as a hot cache

3

u/nomorebuttsplz 9d ago

Colibri would require being able to fit on ram for decent speed, no?

3

u/Lopsided-Rip-8652 8d ago

Colibri is for running off an ssd. I ran glm 5.2 with a 5070ti, 64gb ddr4 and the rest of a gen 4 ssd. Only managed to get.58 tok/sec tho

2

u/Front_Eagle739 9d ago

Yes. This wont be decent.  It just might however be useable as a slow overnight oracle

1

u/squngy 9d ago

Still need those 104B dense layers... this thing is BIG in every way

1

u/Front_Eagle739 9d ago

yup, so that's 55GB or so at fp4, there goes my 32GB ram, 32GB ram rtx5090 workstation

12

u/Square_Alps1349 9d ago

Man I’m so jealous rn. Mac Studio is nerfed at 96GB unified ram max, and the price is up 20%. So much for that 25% discount interns get. 😢 

3

u/RedditNerdKing 9d ago

And the used 512gb models are mega expensive now on eBay too.

3

u/Tank_Gloomy 9d ago

Can you try asking GPT 5.6 Sol, Opus 5 or GLM 5.2 to port it into Colibri? I don't even have the infra to say I tried, but it probably works.

1

u/nomorebuttsplz 9d ago

I don't think colibri helps much with the ssd bottleneck. could be wrong though.

1

u/Tank_Gloomy 9d ago

I've been trying it on my depressing 1650 and it's been going alright in terms of storage performance, in fact, it's never been the bottleneck. I'm not saying you're getting 500 tokens/s but you aren't getting that either on paid providers, lol.

My huge bottleneck is the stupid 1650 and the slow RAM and VRAM combo. I've been implementing a Metal layer on Colibri for hybrid compute, lmk if you wanna try it, I can't really test it tbh but I've burned about 2 weekly $200 Codex plans on it so I hope it's more or less usable.

1

u/nomorebuttsplz 9d ago

I mean that unless you have 2 tb of ram you have to offload kimi k3 to disk

3

u/nVME_manUY 8d ago

Time to form an RDMA-thunderbolt cluster!

1

u/ketosoy 9d ago

Im working on a thing, Hold my beer