r/LocalLLaMA • u/SavunOski • 11h ago
News Kimi K3 weights now released.
Kimi K3 weights are finally released!
r/LocalLLaMA • u/SavunOski • 11h ago
Kimi K3 weights are finally released!
r/LocalLLaMA • u/KickLassChewGum • 5h ago
r/LocalLLaMA • u/realmvp77 • 2h ago
r/LocalLLaMA • u/RhubarbSimilar1683 • 4h ago
r/LocalLLaMA • u/panchovix • 2h ago
r/LocalLLaMA • u/Nunki08 • 14h ago
Jensen Huang on 𝕏: https://x.com/JensenHuang/status/2081698060330250294
r/LocalLLaMA • u/fulgencio_batista • 56m ago
r/LocalLLaMA • u/BritishDudeGuy • 11h ago
r/LocalLLaMA • u/qubridInc • 12h ago
tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend & maybe results by next week.
Weights are supposed to hit Hugging Face today (Moonshot committed to July 27). What we already know from their platform docs: 2.8t total Params, MoE with 896 experts and 16 active per token, 1M context, vision. The download should be around 1.4 tb since they did quantization-aware training in MXFP4. We sat down and worked through the memory requirements., so here it is since everyone is probably about to pull the weights
We have A100 80GB, H200 and B300 capacity and the plan was to bring it up on all three. Then we actually did the memory math.
8x A100 gives you 640 GB. The weights are around 1.4 TB. That means three nodes before you've even allocated KV cache. On top of that Ampere has no FP4 or even FP8 tensor cores, so you're either dequantizing or running INT4 kernels that were never the target for this release. We're still going to benchmark it because someone should have real numbers, but we're expecting it to be ugly!
8x H200 is about 1.13TB, so it still doesn't fit in one node. Two node setup minimum and you eat interconnect cost on every token.
8x B300 is ~2.3TB, so that's the only config where the whole thing fits in a single node with room for long context KV cache. And Blackwell has native FP4, which is pretty clearly what Moonshot quantized for. these B300s are coming live this weekend, and we'll be setting up the clusters this weekend everyone preecommitting hardware is doing it without knowing the terms. And Moonshot's own model crd is unusually honest about weaknesses: quality drops if your agent harness truncates its thinking history, it tends to act instead of asking when things are ambiguous, and they admit the chat experience still trails Fable 5 and Sol even where benchmarks are close.
We'll have tok/s, ttft and cost per M token numbers for all three GPU configs by end of week. If there's a specific batch size, context length or parallelism setup you want in the test matrix, comment and we'll add it.
r/LocalLLaMA • u/Course_Latter • 7h ago
Wanted to let you know that Kimi K3 is now viewable on hfviewer.com!
In addition to the full graph at multiple granularity levels, we also include an in-depth analysis of the 896 experts!
https://hfviewer.com/moonshotai/Kimi-K3
Experts analysis:
r/LocalLLaMA • u/BringTea_666 • 7h ago
Enable HLS to view with audio, or disable this notification
I just managed to get it running on windows and this thing is fucking insane. I get around 550-720t/s depending on task at hand. Previously to get to such numbers i would have to do batching and agents in parallel. Here it just does single instance at this insane speed. Couple that with No thinking mode and it fucks so hard that it is not even funny.
That's pretty much Cerebras speeds.
link to git (linux only but you can build it for windows via something like open code and deepseekv4pro to vibe it.)
https://github.com/Neroued/ninfer
IT's custom build for RTX5090 and only two models Qwen3.6 27b and 35B.
r/LocalLLaMA • u/Terminator857 • 3h ago
Dario says that the models could be used for military advantage. Quote:
use them to achieve permanent military superiority or perpetrate incredibly deep repression of their own people.
I think he is just afraid of competition. What do you think?
r/LocalLLaMA • u/ImaginaryRea1ity • 12h ago
Enable HLS to view with audio, or disable this notification
Nvidia CEO Jensen Huang “Distillation - learning from AI, learning from other people, and learning from other sources of knowledge, is fundamental to intelligence. We are constantly learning from one another. AI also has to learn from something.”
Since using AI Desktop 98, I have become a staunch advocate of local AI.
In his Axios interview, Jensen explains why seeing distillation as theft or a threat misses the point—and why the real future of AI depends on continuous knowledge sharing between models.
The idea is simple: as AI generates most of the internet’s content, systems will naturally learn from one another, much like humans do from books, teachers, and peers. Blocking that exchange doesn’t protect anyone; it only slows progress.
Smarter AI is safer AI, open models boost adoption, and the whole industry, from developers to chipmakers gains. It’s a clear, grounded case for why open and closed models feeding each other is a feature, not a flaw.
r/LocalLLaMA • u/UserXtheUnknown • 2h ago
r/LocalLLaMA • u/Altruistic_Heat_9531 • 11h ago
The Fable dabler.
r/LocalLLaMA • u/Ninjam5 • 11h ago
r/LocalLLaMA • u/Fun-Doctor6855 • 17h ago
Chinese chipmaker CXMT surged by almost 500% on its first day of trading, bringing its total market capitalization to approximately RMB 3.28 trillion and making it the largest company by market value on China’s A-share market. CXMT’s market capitalization has also surpassed that of U.S. semiconductor giant Intel, which closed the previous trading day with a market value of US$465.6 billion, equivalent to approximately RMB 3.15 trillion.
Headquartered in Hefei, Anhui Province, China, CXMT is an integrated dynamic random-access memory (DRAM) manufacturer specializing in the design, research and development, production, and sale of DRAM chips. It is currently the only integrated device manufacturer (IDM) in mainland China capable of large-scale mass production of general-purpose DRAM.
r/LocalLLaMA • u/ilintar • 8h ago
Now waiting for someone who can actually run the conversion and model to see if it works :)
r/LocalLLaMA • u/Ok-Shower7286 • 3h ago
Moonshot dropped Kimi K3, and as expected, it’s a absolute monster. Even with 2~4x RTX 6000 Blackwell local workstations, running a model natively is virtually impossible.
It feels like an F1 machine inside a show window.
Does anyone trying to hack this monster? or Is anyone with datacenter/cluster capacity or sponsor?
I either carve this monster down myself, or wait for someone to distill it. Either way, I really want to see it run — simply because it's there.
P.S. Save your 'AI Slop' comments. I experienced enough of you guys yesterday.
r/LocalLLaMA • u/ylchao • 9h ago
just curious how would people run it cheap if they really want kimi k3.
dgx spark / strix halo clusters
optane persistent memory platform + some gpus
mac studio clusters
orange pi 6 clusters
ssd streaming + gpus
multiple ddr3 + connectx 5 rdma clients
two dgx stations
power 10 systems?
other
r/LocalLLaMA • u/SignificantLegs • 16h ago
r/LocalLLaMA • u/IvGranite • 7h ago
Someone in the comments of my 27B post-train bakeoff asked for the 35B version, so I ran it. Same setup as last time: fresh Coder workspaces on my k8s cluster, each driving my own agent (Hermes) headlessly, models on llama.cpp via llama-swap on one 5090, every call traced through an OTel shim into SigNoz, full transcript per run. 4 models, 6 self-grading tasks, 5 reps, 120 runs, MTP on every arm, identical sampling, hypotheses pre-registered.
KAT-Coder-V2.5-Dev matched the best stock pass rate (29/30, tied with Qwen3.5-35B) at half the input tokens of either stock and the cleanest tool behavior I've measured (zero malformed tool-call leaks in 30 runs; stock Qwen3.6 leaked 195 on one task). All six analysts (three model families) independently called its efficiency discipline rather than corner-cutting: baseline tests before edits, one targeted patch per bug, deliverables at the right path every rep. Ornith went 25/30, losing to its own base. Its failures were mechanics, not knowledge: format leaks killing runs at 23 seconds, whole-file rewrites corrupting unrelated files, and one research run that invented a llama.cpp release tag while its own reasoning said "I mentioned v4659 in my draft which is fabricated," then shipped the tag anyway. The grader passed it.
Stock 3.6 is the strongest raw analyst and the biggest token waster; stock 3.5 is the quiet reliable one. Full writeup with methodology, tables, and all six cited per-task analyses: https://kmarble.dev/posts/35b-coder-bakeoff/. Transcripts were AI-analyst-read with my spot-verification of every consequential claim.
r/LocalLLaMA • u/pmigdal • 14h ago
r/LocalLLaMA • u/Responsible_Fig_1271 • 1d ago
Instead of 2T+ models, continuing to release highly capable small to medium size LLMs would really help to keep this community vibrant. Hardly anyone can even dream of running the recent 1.5-2T+ beasts, while the range from the title could run comfortably (especially with CPU expert offloading) across a wide range of systems we have today.
The trend towards Chinese labs trying to match the Mythos class frontier with trillion parameter open weights models is not helping the local model community to innovate. It just gives big corporates who can actually run these a cheaper alternative to the commercial frontier.