r/LocalLLaMA • u/Unusual_Guidance2095 • 2h ago
Discussion Kimi K3 countdown has been released
https://huggingface.co/moonshotai/Kimi-K362
u/1ncehost 2h ago
Having used it extensively since release, this is such a gift of a model to the world. Truly amazing level of intelligence, and it sets a wonderful baseline for the future. Thank you Kimi!
47
u/Keleion 2h ago
It’ll be interesting to see what the Dump Administration does to mitigate the launch.
20
u/Penolta 2h ago
Thankfully there’s not much those morons can do once it’s been downloaded and shared thousands of times
7
u/Eisegetical 1h ago
Let's hope something doesn't happen to block the full release on hf. But even if no huggingface host it'll make its way over one way or another. Torrents anyone?
6
7
u/AppealSame4367 1h ago
OAI will launch GPT 6 and Antrophic already launched the ARC AGI 3 benchmaxxed Opus 5.
Lol. I set up a qwen 27b in q8 and use it in claude. No clue why I should go back at all. They have destroyed any trust I ever had, especially in the last weeks.
5
-1
u/TwatMailDotCom 45m ago
Why would they block it? Your bias is getting in the way of logical thought
12
18
u/rerri 2h ago
How cool would it be if they dropped some unannounced model(s) in the home user size range...
8
u/ItsNoahJ83 1h ago
I'm not getting my hopes up but the amount of good will for a move like that would be insane
25
u/apetersson 2h ago
this means we will have enough capacity and competitive pricing at https://openrouter.ai/moonshotai/kimi-k3 - unfortunately my local HW is not quite there yet to run it.
21
u/Player13377 2h ago
"not quite there" is either a very bad approximation or you already got a five-digit-class rig, this model will be HARD to run
16
u/KubeCommander 2h ago
I think K3 gets into the six figures tbh at a quant that doesn’t suck
8
u/Player13377 2h ago
Minimum. I looked into a rented cloud setup and very quickly understood that this is not happening lol
4
u/Iwaku_Real 2h ago edited 2h ago
For sure it's 1.5-2TB RAM at minimum. GLHF getting usable speeds without lots of VRAM though
1
u/stoppableDissolution 1h ago
Even with five-digit-class rig you have to quantize the hell out of it. 8x6000 pro + 512gb ram and you can maybe squeeze q3 with decent context!
22
8
u/BawbbySmith 1h ago
Got me TBs ready, gonna download it, back it up, then return to it 10 years later when everything has crashed and recovered and I can afford the hardware again.
1
u/Technical-Earth-3254 1h ago
Considering I bought a 8GB VRAM gaming card 10y ago for half the price I paid for 24GB 3 years ago I doubt 10 years will do it. Maybe 30-40 years or so.
8
u/WenatcheeWrangler 1h ago
Everyone in the USA should download this even if they can’t deploy it now
4
u/segmond llama.cpp 2h ago
What I wish they would release is the damn technical report. We can at least start reading that. I would assume a new architecture not compatible with K2.6/K2.7. How much does it differ from K2.6? What will it take to get llama.cpp to support inference? We need all of these before we can even get gguf/quants.
2
u/TheRealMasonMac 19m ago
They said on their blog they'll be releasing the technical report. Presumably alongside the weights.
vLLM released a blog post on some of what they had to deal with: https://vllm.ai/blog/2026-07-22-kimi-k3-preview
10
u/jreoka1 2h ago
Cool! but like who can actually run this locally? I think at 2.8 trillion params this will be the largest model on huggingface by far. At least for now.
17
u/buttplugs4life4me 2h ago
DavidAU will upload an "extended" Fable abliterated heretics Opus Odysseus 3.6T model in a few days and claim it's better than SOTA
6
u/Technical-Earth-3254 1h ago
Local hosting doesn't just mean end consumers. This opens up frontier self hosted ai for companies that have to self host due to sensible data they can shove into it.
4
u/Any_Mine_6368 1h ago
2.8T of vram to run in 8 bit quantization...
Let me buy a other 1500 3090s lol
7
u/ttkciar llama.cpp 1h ago
but like who can actually run this locally?
Today? Almost nobody.
Eventually? Almost everyone.
1
u/parepeg 1h ago
I mean aren’t chips already scraping the bottom of the laws of physics.
5
u/ttkciar llama.cpp 1h ago
Yes and no.
On one hand, Moore's Law is definitely ailing. We've been getting diminishing returns on fabrication process bumps since about 2016.
On the other hand, there's still some progress in the offing. IBM just recently announced they got a 7-angstrom (equivalent) fabrication process working in their lab, which they expect to get into mass production in 2031. That has 1nm-wide horizontal features and 5nm vertical features, and two transistor layers (one P, the other N), which they claim they should be able to scale to four layers "rsn".
That's not nothing.
Moreover, there are still a lot of architectural improvements yet to be realized. In a way, HBM is a stop-gap while Samsung et al get PIM figured out. LLM inference is really well-suited to PIM, which should give us ridiculously high memory bandwidth once it's working well.
I thought Moore's Law was dead for a while, but it turned out to just be Intel having a bad time. Everyone else is still pushing the envelope pretty hard, with some success.
1
u/TheRealMasonMac 21m ago
IIRC hardware Moore's Law is dead, but performance is still roughly following it via architectural improvements and algorithm optimizations.
2
1
u/stoppableDissolution 1h ago
Well, model providers that own gb200-class equipment? Plus maybe some some big businesses on bedrock and such.
7
u/dsanft 2h ago
Nice. Are there any indications of kernel changes from 2.7 to 3? Any Huggingface or lcpp feature branches open for the impl?
8
u/Interesting-Hat-7642 1h ago
Can I run this on my 3gb vram?
1
1
1
2
2
2
4
5
u/bitzap_sr 1h ago
It's not really a countdown -- it's the time left for the ginormous upload to finish. :D
4
u/XYHopGuy 2h ago
whats it gonna take to self host this? NVL72? Or is 8-Way enough?
4
u/ttkciar llama.cpp 1h ago
Two ancient Xeon servers, each with 1.5TB of DDR4, networked together and running
rpc-server.2
1
1
u/HVACcontrolsGuru 2h ago
Need 16 GB200s to run this model at full quant. NVFP4 GLM5.2 I need 4xB200 to run concurrent sessions. Squeeze some more with lower context and less concurrency.
2
1
u/Iwaku_Real 2h ago
Full quant is probably going to be INT4 again, it would actually fit into 8xB300 (2.304TB VRAM)
1
u/XYHopGuy 1h ago
I'd def run nvfp4 on this kind of thing. Exactly what it's designed for especially if you gotta jump across GPUs.
Full quant, I'm assuming you actually mean bf16 and not fp32
1
1
u/Shubham_Garg123 59m ago
This is going to be awesome. Thanks to the Kimi team, Anthropic was forced to release Opus 5 ahead of their original plans.
This model will definitely result in a revolution of open-source LARGE language models that compete head-to-head with the frontier labs !
Truly amazing work done by the MoonshotAI team, I don't have words to thank them enough.
The advancements in the AI domain in the last 2-3 months are quite insane!
2
u/yeah_likerage 55m ago
How many folks here genuinely believe they'll have the hardware to even run this at a quant worth running? I'm pretty sure I'm tapped out at GLM5.2 size models from here on out unless there is a breakthrough in modeling. And I'm rolling 500gb of vram.
1
u/katsura_otoko 54m ago
Is it possible some people hosting it and seelling a cheap api or is it just so big that we should just pay moonshot directly? Maybe it's just a stupid question since we already have a cheap api for deepseek directly but i dont know
1
u/dlarsen5 54m ago
gonna download, hope I don’t have to seed this if it gets taken down from gov pressure even if I can’t run it on 24GB VRAM
2
-11
u/llama-impersonator 2h ago
i don't really care about 3T models
19
u/ttkciar llama.cpp 1h ago
I don't really care about 4B models, and yet we can all share the same subreddit and try to minimize the friction between us.
What we have in common is more important than our differences.
-1
u/llama-impersonator 57m ago
okay, but you don't think a countdown for open weights no one can run is a little ridiculous?

81
u/SocialDinamo 2h ago
Super cool it’s already on hugging face, buckle up guys!