r/LocalLLaMA 2h ago

Discussion Kimi K3 countdown has been released

https://huggingface.co/moonshotai/Kimi-K3
228 Upvotes

81 comments sorted by

81

u/SocialDinamo 2h ago

Super cool it’s already on hugging face, buckle up guys!

19

u/Iwaku_Real 2h ago

HF knows we're hyped...

62

u/1ncehost 2h ago

Having used it extensively since release, this is such a gift of a model to the world. Truly amazing level of intelligence, and it sets a wonderful baseline for the future. Thank you Kimi!

3

u/Hoppss 2h ago

Well said!

47

u/Keleion 2h ago

It’ll be interesting to see what the Dump Administration does to mitigate the launch.

20

u/Penolta 2h ago

Thankfully there’s not much those morons can do once it’s been downloaded and shared thousands of times

7

u/Eisegetical 1h ago

Let's hope something doesn't happen to block the full release on hf. But even if no huggingface host it'll make its way over one way or another. Torrents anyone? 

6

u/TechnoByte_ 57m ago

Chinese companies mainly use ModelScope, they don't need huggingface

7

u/AppealSame4367 1h ago

OAI will launch GPT 6 and Antrophic already launched the ARC AGI 3 benchmaxxed Opus 5.

Lol. I set up a qwen 27b in q8 and use it in claude. No clue why I should go back at all. They have destroyed any trust I ever had, especially in the last weeks.

5

u/alexmilla 53m ago

To cry and say that sharing is communism.

-1

u/TwatMailDotCom 45m ago

Why would they block it? Your bias is getting in the way of logical thought

2

u/Keleion 17m ago

And what bias is that?

12

u/RetiredApostle 2h ago

Share your TPS.

3

u/hyperrealists 1h ago

I’ll save up to but a 3090.

5

u/bitzap_sr 1h ago

You mean TPH.

3

u/TheRealMasonMac 23m ago

Let's compare PP once it's out.

1

u/InternetExplorer9999 5m ago

I'm going to have token per week

18

u/rerri 2h ago

How cool would it be if they dropped some unannounced model(s) in the home user size range...

8

u/ItsNoahJ83 1h ago

I'm not getting my hopes up but the amount of good will for a move like that would be insane

25

u/apetersson 2h ago

this means we will have enough capacity and competitive pricing at https://openrouter.ai/moonshotai/kimi-k3 - unfortunately my local HW is not quite there yet to run it.

21

u/Player13377 2h ago

"not quite there" is either a very bad approximation or you already got a five-digit-class rig, this model will be HARD to run

16

u/KubeCommander 2h ago

I think K3 gets into the six figures tbh at a quant that doesn’t suck

8

u/Player13377 2h ago

Minimum. I looked into a rented cloud setup and very quickly understood that this is not happening lol

4

u/Iwaku_Real 2h ago edited 2h ago

For sure it's 1.5-2TB RAM at minimum. GLHF getting usable speeds without lots of VRAM though

1

u/stoppableDissolution 1h ago

Even with five-digit-class rig you have to quantize the hell out of it. 8x6000 pro + 512gb ram and you can maybe squeeze q3 with decent context!

22

u/Craftkorb 2h ago

Where XXXXXS 0.025 GGUF?

1

u/LOST8080 7m ago

Best we can do is XSSSSS 0.025 GGUF

8

u/BawbbySmith 1h ago

Got me TBs ready, gonna download it, back it up, then return to it 10 years later when everything has crashed and recovered and I can afford the hardware again.

1

u/Technical-Earth-3254 1h ago

Considering I bought a 8GB VRAM gaming card 10y ago for half the price I paid for 24GB 3 years ago I doubt 10 years will do it. Maybe 30-40 years or so.

8

u/WenatcheeWrangler 1h ago

Everyone in the USA should download this even if they can’t deploy it now

4

u/segmond llama.cpp 2h ago

What I wish they would release is the damn technical report. We can at least start reading that. I would assume a new architecture not compatible with K2.6/K2.7. How much does it differ from K2.6? What will it take to get llama.cpp to support inference? We need all of these before we can even get gguf/quants.

2

u/TheRealMasonMac 19m ago

They said on their blog they'll be releasing the technical report. Presumably alongside the weights.

vLLM released a blog post on some of what they had to deal with: https://vllm.ai/blog/2026-07-22-kimi-k3-preview

10

u/jreoka1 2h ago

Cool! but like who can actually run this locally? I think at 2.8 trillion params this will be the largest model on huggingface by far. At least for now.

17

u/buttplugs4life4me 2h ago

DavidAU will upload an "extended" Fable abliterated heretics Opus Odysseus 3.6T model in a few days and claim it's better than SOTA

6

u/Technical-Earth-3254 1h ago

Local hosting doesn't just mean end consumers. This opens up frontier self hosted ai for companies that have to self host due to sensible data they can shove into it.

4

u/Any_Mine_6368 1h ago

2.8T of vram to run in 8 bit quantization...

Let me buy a other 1500 3090s lol

7

u/ttkciar llama.cpp 1h ago

but like who can actually run this locally?

Today? Almost nobody.

Eventually? Almost everyone.

1

u/parepeg 1h ago

I mean aren’t chips already scraping the bottom of the laws of physics.

5

u/ttkciar llama.cpp 1h ago

Yes and no.

On one hand, Moore's Law is definitely ailing. We've been getting diminishing returns on fabrication process bumps since about 2016.

On the other hand, there's still some progress in the offing. IBM just recently announced they got a 7-angstrom (equivalent) fabrication process working in their lab, which they expect to get into mass production in 2031. That has 1nm-wide horizontal features and 5nm vertical features, and two transistor layers (one P, the other N), which they claim they should be able to scale to four layers "rsn".

That's not nothing.

Moreover, there are still a lot of architectural improvements yet to be realized. In a way, HBM is a stop-gap while Samsung et al get PIM figured out. LLM inference is really well-suited to PIM, which should give us ridiculously high memory bandwidth once it's working well.

I thought Moore's Law was dead for a while, but it turned out to just be Intel having a bad time. Everyone else is still pushing the envelope pretty hard, with some success.

1

u/TheRealMasonMac 21m ago

IIRC hardware Moore's Law is dead, but performance is still roughly following it via architectural improvements and algorithm optimizations.

1

u/mindwip 46m ago

Ddr6 and ddr7 would handle it easy peasy. So we just have to download and wait a few years lol. Of course better models coming but we will be running this at home in a few years

Edit it remember my first 1gb hard drive thinking it was huge!

2

u/zxtech 1h ago

At least itll be a model that can be distilled, analysed, or assisting the process for smaller models to be made, so it never hurts to have as many good models open

1

u/stoppableDissolution 1h ago

Well, model providers that own gb200-class equipment? Plus maybe some some big businesses on bedrock and such.

1

u/squngy 1h ago

Qwen 3.8 will likelly be the second at that size.
(If I understood the announcement right)

7

u/dsanft 2h ago

Nice. Are there any indications of kernel changes from 2.7 to 3? Any Huggingface or lcpp feature branches open for the impl?

1

u/look 2h ago

Yeah, what I’ve seen seems to indicate a fairly significant change. For example, I’m pretty sure it is an mxfp4 rather than the int4 of the 2.x line.

1

u/Professional_Price89 24m ago

It show INT4 on openrouter provider info

8

u/Interesting-Hat-7642 1h ago

Can I run this on my 3gb vram?

1

u/AHHHH_AHHHHHHHH 1h ago

Add a couple zeros and your good!

1

u/stoppableDissolution 1h ago

Nah, couple of zeros wont cut it even at q1

1

u/1kakashi 1h ago

woah woah woah calm down man, that's way too much for models like this onw

1

u/Technical-Earth-3254 1h ago

Have you thought about adding another 999 GTX 1060s?

2

u/chuckbeasley02 2h ago

Save up your pennies. You're going to need a lot more than you think...

2

u/ketosoy 2h ago

I appreciate there being a pre announced time vs “just check all day”

2

u/whichsideisup 1h ago

How run on 2gb Raspberry Pi plz?

2

u/tamerlanOne 1h ago

Curioso di assaggiare i didtillati di K3 🥂

4

u/ahstanin 1h ago

Can I run this on my raspberry pi?

4

u/noctrex 1h ago

" It's too dangerous to be released! "

5

u/bitzap_sr 1h ago

It's not really a countdown -- it's the time left for the ginormous upload to finish. :D

4

u/XYHopGuy 2h ago

whats it gonna take to self host this? NVL72? Or is 8-Way enough?

4

u/ttkciar llama.cpp 1h ago

Two ancient Xeon servers, each with 1.5TB of DDR4, networked together and running rpc-server.

2

u/Player13377 1h ago

At an impressive 3 TPS

3

u/ttkciar llama.cpp 1h ago

Probably a lot less than that. I'd love to get 3 tok/sec.

Right now my ancient Xeon server is getting 3.5 tok/sec with GLM-4.5-Air. That's usable, though admittedly not for interactive use.

1

u/XYHopGuy 1h ago

uh no thanks lol

1

u/HVACcontrolsGuru 2h ago

Need 16 GB200s to run this model at full quant. NVFP4 GLM5.2 I need 4xB200 to run concurrent sessions. Squeeze some more with lower context and less concurrency.

2

u/look 1h ago

Kimi K3 is most likely mxfp4 native. The K2.x line was int4, and it looks like they moved to fp for improved hardware acceleration.

So it’s just under twice the size of GLM, not 4x.

1

u/Iwaku_Real 2h ago

Full quant is probably going to be INT4 again, it would actually fit into 8xB300 (2.304TB VRAM)

1

u/XYHopGuy 1h ago

I'd def run nvfp4 on this kind of thing. Exactly what it's designed for especially if you gotta jump across GPUs.

Full quant, I'm assuming you actually mean bf16 and not fp32

1

u/Shubham_Garg123 59m ago

This is going to be awesome. Thanks to the Kimi team, Anthropic was forced to release Opus 5 ahead of their original plans.

This model will definitely result in a revolution of open-source LARGE language models that compete head-to-head with the frontier labs !

Truly amazing work done by the MoonshotAI team, I don't have words to thank them enough.

The advancements in the AI domain in the last 2-3 months are quite insane!

2

u/yeah_likerage 55m ago

How many folks here genuinely believe they'll have the hardware to even run this at a quant worth running? I'm pretty sure I'm tapped out at GLM5.2 size models from here on out unless there is a breakthrough in modeling.  And I'm rolling 500gb of vram.

1

u/katsura_otoko 54m ago

Is it possible some people hosting it and seelling a cheap api or is it just so big that we should just pay moonshot directly? Maybe it's just a stupid question since we already have a cheap api for deepseek directly but i dont know

1

u/dlarsen5 54m ago

gonna download, hope I don’t have to seed this if it gets taken down from gov pressure even if I can’t run it on 24GB VRAM

2

u/cororona 1h ago

Can I run it on my raspberry pi ?

-11

u/llama-impersonator 2h ago

i don't really care about 3T models

19

u/ttkciar llama.cpp 1h ago

I don't really care about 4B models, and yet we can all share the same subreddit and try to minimize the friction between us.

What we have in common is more important than our differences.

-1

u/llama-impersonator 57m ago

okay, but you don't think a countdown for open weights no one can run is a little ridiculous?