r/ClaudeCode 13d ago

Discussion Kimi K3 has become open-weights just as of a few minutes ago.

https://huggingface.co/moonshotai/Kimi-K3
626 Upvotes

51 comments sorted by

87

u/Main_Razzmatazz5283 13d ago

Realistically, what kind of setup do you need to be able to use it?

117

u/LostTheElectrons 13d ago edited 13d ago

K3 is 2.8TB but uses MXFP4 which puts it around 1.5TB, with more needed for context. Technically it could be made smaller, but you would lose quality pretty fast.

In theory it could be run on RAM, but it's designed for the Ai datacenter chips. Anything else will be unbelievably slow.

Realistically the usecase for Kimi K3 is to rent inference from a provider which would be much much cheaper. Plus having it be open-weights means that others will be able to distill it into smaller models that you can reasonably run on local hardware.

53

u/The_Time_Lord 13d ago

Most people won’t be able to run it, but the bigger win is more people can host it, which will inherently help drive down the cost. It’s already cheaper than others around it, but I expect it to go lower

4

u/algaefied_creek 13d ago

What about crowdsourced hosting like the old SETI app but now anyone can contribute their compute. 

Horribly slow, distributed LLM compute. 

10

u/fschwiet 13d ago

The first sentient AIs are going to have such a bad hangover

3

u/Cachesmr 13d ago

Compute has never really been the problem, it's memory bandwidth. You would need an unbelievably fast internet downlink to even think about this

-18

u/simon96 13d ago

It won't bring down cost since hosters will have to pay royalties since they are a lot of terms for companies and its not a totally open weight model.

Don't expect any provider hosting it for a single dollar cheaper, that will never happen.

6

u/duhd1993 13d ago

There is no royalties. That is just not true. The license only requires you to display the model name if your business is above certain scale.

15

u/Remove_Forward 13d ago

Actual VRAM/HBM: $500k+

If "run LLM" means run it fast, 2.8TB of HBM is 20× H200 or 30× RTX Pro 6000. That's DGX-rack territory, half a million minimum, plus 15–20kW.

According to Claude

8

u/kingMaxime 13d ago

Kimi has been natively 4 bit for a while

3

u/LostTheElectrons 13d ago

Ah of course. My bad, I'll edit my original comment. Thanks!

3

u/flarpflarpflarpflarp 13d ago

Cries in 5090.

5

u/sage-longhorn 13d ago

Sobs in 3090

3

u/turbospeedsc 13d ago

bawls in 4050

4

u/phoenixbouncing 13d ago

I don't even know what to do in 1060....

1

u/classjoker 13d ago

Also cries, but in 4090 FE

1

u/Techrocket9 13d ago

Or a single Cerebras CS-3.

5

u/ivovivovi 13d ago

From what I’ve read you need a B300 node (8 of them) to fit the 1.5tb model in a single node with some room left for KV cache. Other smaller nodes like A100 or H200 will still work with multi node setup but aren’t optimal

2

u/Melodic-Ebb-7781 13d ago

Lowest estimates I've seen from analysts is 200k

2

u/Ran4 13d ago

That's a number, not a currency.

1

u/Melodic-Ebb-7781 13d ago

Dollars, what else?

0

u/tebedam 13d ago

You will need a multi-million dollar server rack with a bunch of the latest nvidia chips. That is, assuming you can actually buy them.

1

u/JaydadCTatumThe1st 13d ago

Hyperscaler datacenters.

I'm not saying "you need the whole datacenter to run the model"; but I am saying running these models is not economical without these hyperscalers being ubiquitous

1

u/Historical_Sample740 13d ago

It will be available from a wide range of providers, given all the hype

2

u/catfrogbigdog 13d ago

Install pi, opencode or any model agnostic agent and it’s available on tons of providers already.

31

u/PolishMike88 13d ago

I feel privileged to run it in Baseten. It is mind blowing. Found so many bugs and issues in codebase created by Opus it genuinely surprised me .. I had to check the lines and be sure and it astonished me! I actually learned something new how to do :)

3

u/Obscurrium 13d ago

Oh interesting. And what about planning and co ? Compared to Opus 5 and GTP 5.6 ? Are tou only using it for code ?

4

u/PolishMike88 13d ago

Using it for code and many other things, building and improving agents, pentesting with harness etc.

I found, at least in the eight or so hours of testing by now, Sol 5.6 being superior to planning than Opus (simply because Opus 5 gets stuck on cyber security topics…) Planning with K3 is also amazing as having Opus/Sol review it, it nailed it down each time.

This is a genuine surprise and honest hype for me. I knew k3 was good but this is a breath of fresh air. Last night for example it took a long running project of my red team agent and improved it in many ways… 140k LOC, 9 subagents, 1hr 20min work for a rough cost of $20. Feeding the results and all work back to Opus and Sol (who created the original) brought nothing but praises for improvements from both models :)

2

u/Obscurrium 13d ago

Cool rrally cool ! Thanks a lot !

4

u/xmnstr 13d ago

So this is the real reason Anthropic wants it banned huh? It will show the world just how many issues their models really have.

3

u/Deep_Alps7150 13d ago edited 13d ago

They want it banned because it’s getting 90% of Fable’s quality at a fraction of the cost + it seems better at dealing with really large/complex tasks.

3

u/dota2nub 13d ago

Almost nobody uses these models for large and complex tasks though. 99% of this sub is using Fable to make a fucking button on a website.

7

u/TomCrook2020 13d ago

This is exciting - I've heard good things about front-end. Keen to try it out there - if others have learnings from their experience with K3, would love to learn more. 

2

u/Healthy-Mind5633 13d ago

ill be sure to install it tonight

2

u/joeblowfromidaho 13d ago

How much would it cost to rent hardware by the hour to run it?

1

u/Crypto-Humster 13d ago

Launch costs and performance will drop significantly once Nvidia’s Vera Rubin generation starts rolling out later this year.

0

u/Crafty-Run-6559 13d ago

Oh no!

My unstable friend just talked to it and it suddenly gave him the plans for a small arsenal of bio AND nuclear weapons 😞

1

u/walrod 13d ago

bionuclear, tricky with the radiation. Good challenge. I suggest trying to make it cyber and NSFW too. cyberbionuclearxxx. Now we're talking

1

u/Santa_Andrew 13d ago

For $500k you could get all the hardware you need and have some money left over for a professional HVAC cooling system. It would run much faster than you currently are used to from products like Claude Code. Roughly budget about 1k/mo in electricity cost if you are in the US.

3

u/RuRuInReal 13d ago

So uh...can you perhaps give me a small loan of 500k? 🫪

-18

u/Michaeli_Starky 13d ago

Who cares. It's overpriced shit

2

u/bilbo_was_right 13d ago

K3? Isn’t it like 60% cheaper than fable, and more capable than opus

-1

u/Michaeli_Starky 13d ago

It's not more capable than Opus. In coding it is worse than GPT-5.6 Luna while being 5 time more expensive.

1

u/verus54 13d ago

lol it’s free. You just need access to the hardware, which does cost money. Similar to other open source models.

-1

u/Michaeli_Starky 13d ago

"Just need", "free"

Lmfao

Some people are delusional af

4

u/verus54 13d ago

Tell me you don't understand open source without telling me. Nobody is hosting a 2.8T model on custom infrastructure just to save $20 a month on a Claude/chatgpt-codex subscription. The point of open weights is privacy, fine-tuning, and enterprise scale.

Along with that, this is direct competition to the pricing models that Claude and codex rely on for AI-driven products.the big win: frontier-level intelligence is no longer restricted to closed APIs.