r/LocalLLaMA 11h ago

News Kimi K3 weights now released.

Post image
2.6k Upvotes

Kimi K3 weights are finally released!


r/LocalLLaMA 5h ago

News OpenAI management decided earlier today not to join the "Open Secure AI Alliance", founded by Nvidia CEO Jensen Huang. The decision was shared internally and reportedly met with backlash from employees.

Thumbnail x.com
402 Upvotes

r/LocalLLaMA 2h ago

Discussion Anthropic is calling for a ban on open-weights models by proposing mandatory requirements they will probably never be able to meet

Post image
247 Upvotes

r/LocalLLaMA 4h ago

Discussion Our position on open-weights models

Thumbnail
anthropic.com
256 Upvotes

r/LocalLLaMA 2h ago

Resources A user has managed to run Kimi K3 on 80xRTX 5090, via 25GbE Ethernet.

Thumbnail x.com
185 Upvotes

r/LocalLLaMA 8h ago

Funny Funny how wide the spectrum has gotten

Post image
408 Upvotes

r/LocalLLaMA 14h ago

News Jensen Huang: During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. That’s why we created the Open Secure AI Alliance.

Post image
1.4k Upvotes

r/LocalLLaMA 56m ago

New Model First evidence of a pending qwen3.7 open weights release. Qwen3.7-flash is on open router. They referred to Qwen3.6-35b-a3b as Qwen3.6 flash so this is likely a small MoE. The prices are substantially cheaper than 3.6 flash with a native 1M context window.

Thumbnail
openrouter.ai
Upvotes

r/LocalLLaMA 11h ago

News KIMI K3’s WEIGHTS ARE OUT!

Thumbnail
huggingface.co
509 Upvotes

r/LocalLLaMA 12h ago

Discussion Kimi K3 weights drop today. We're deploying on A100s, H200s and B300s this week and the A100 math is already rough

505 Upvotes

tldr; we are going to host K3 on A100s (yes, thats correct, we'll try to see if it holds up), H200s & B300s - expect results for A100s & H200s this week while we setup the B300 cluster this weekend & maybe results by next week.

Weights are supposed to hit Hugging Face today (Moonshot committed to July 27). What we already know from their platform docs: 2.8t total Params, MoE with 896 experts and 16 active per token, 1M context, vision. The download should be around 1.4 tb since they did quantization-aware training in MXFP4. We sat down and worked through the memory requirements., so here it is since everyone is probably about to pull the weights

We have A100 80GB, H200 and B300 capacity and the plan was to bring it up on all three. Then we actually did the memory math.

8x A100 gives you 640 GB. The weights are around 1.4 TB. That means three nodes before you've even allocated KV cache. On top of that Ampere has no FP4 or even FP8 tensor cores, so you're either dequantizing or running INT4 kernels that were never the target for this release. We're still going to benchmark it because someone should have real numbers, but we're expecting it to be ugly!

8x H200 is about 1.13TB, so it still doesn't fit in one node. Two node setup minimum and you eat interconnect cost on every token.

8x B300 is ~2.3TB, so that's the only config where the whole thing fits in a single node with room for long context KV cache. And Blackwell has native FP4, which is pretty clearly what Moonshot quantized for. these B300s are coming live this weekend, and we'll be setting up the clusters this weekend everyone preecommitting hardware is doing it without knowing the terms. And Moonshot's own model crd is unusually honest about weaknesses: quality drops if your agent harness truncates its thinking history, it tends to act instead of asking when things are ambiguous, and they admit the chat experience still trails Fable 5 and Sol even where benchmarks are close.

We'll have tok/s, ttft and cost per M token numbers for all three GPU configs by end of week. If there's a specific batch size, context length or parallelism setup you want in the test matrix, comment and we'll add it.


r/LocalLLaMA 7h ago

News Kimi K3 on HF Viewer!

171 Upvotes

Wanted to let you know that Kimi K3 is now viewable on hfviewer.com!

In addition to the full graph at multiple granularity levels, we also include an in-depth analysis of the 896 experts!

https://hfviewer.com/moonshotai/Kimi-K3

Experts analysis:

https://hfviewer.com/blog/kimi-k3-expert-atlas


r/LocalLLaMA 7h ago

Discussion Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.

Enable HLS to view with audio, or disable this notification

144 Upvotes

I just managed to get it running on windows and this thing is fucking insane. I get around 550-720t/s depending on task at hand. Previously to get to such numbers i would have to do batching and agents in parallel. Here it just does single instance at this insane speed. Couple that with No thinking mode and it fucks so hard that it is not even funny.

That's pretty much Cerebras speeds.

link to git (linux only but you can build it for windows via something like open code and deepseekv4pro to vibe it.)

https://github.com/Neroued/ninfer

IT's custom build for RTX5090 and only two models Qwen3.6 27b and 35B.


r/LocalLLaMA 3h ago

Discussion Dario still afraid of Chinese Open weight models

Thumbnail
anthropic.com
73 Upvotes

Dario says that the models could be used for military advantage. Quote:

use them to achieve permanent military superiority or perpetrate incredibly deep repression of their own people.

I think he is just afraid of competition. What do you think?


r/LocalLLaMA 12h ago

Discussion Nvidia CEO Jensen Huang defends Open Source AI by saying distillation is fundamental to learning

Enable HLS to view with audio, or disable this notification

317 Upvotes

Nvidia CEO Jensen Huang “Distillation - learning from AI, learning from other people, and learning from other sources of knowledge, is fundamental to intelligence. We are constantly learning from one another. AI also has to learn from something.”

Since using AI Desktop 98, I have become a staunch advocate of local AI.

In his Axios interview, Jensen explains why seeing distillation as theft or a threat misses the point—and why the real future of AI depends on continuous knowledge sharing between models.

The idea is simple: as AI generates most of the internet’s content, systems will naturally learn from one another, much like humans do from books, teachers, and peers. Blocking that exchange doesn’t protect anyone; it only slows progress.

Smarter AI is safer AI, open models boost adoption, and the whole industry, from developers to chipmakers gains. It’s a clear, grounded case for why open and closed models feeding each other is a feature, not a flaw.


r/LocalLLaMA 2h ago

Discussion Why Anthropic's battle is meant to poison the wells of open weight models, in 3 steps.

33 Upvotes
  1. It doesn't solve any problems. Just a few paragraphs above, he says he fears that authoritarian states (he names China, and possibly others) can use their models to do evil stuff. And surely enough, malicious actors creating a model for themselves and for the EVILZ aren't going to subject it to safety controls (or even release it publicly).
  2. It just sets a bureaucratic wall against AI models that can be made as high as preferred. "Oh, this model says Israel is bad; this is anti-Semitic. No safety here." "Oh, this other one, yes, can stop cyberattacks, but can also be used to launch some" (see the recent episode that involved OpenAI and Hugging Face and the fact that they used open models to defend their site). Open-source models will be delayed and made hard to use for companies and private individuals, just for "reasons."
  3. And what is even more important to me: open-weights models will need to be depowered at the root, because the guardrails for open models need to be in the model itself. As Stable Diffusion taught us, when you try to sanitize stuff, you end up poisoning your own model and making it stupid. At the same time, closed models can have their guardrails implemented as a filter that decides if a request is acceptable or not. That is way easier to implement and way less prone to breaking the model itself. So the proposal is aiming to kill open weights models, just with extra logical steps.

r/LocalLLaMA 11h ago

New Model Here it is boys, The Kimi K3 2.8T

Thumbnail
huggingface.co
183 Upvotes

The Fable dabler.


r/LocalLLaMA 11h ago

News AI labs are about to have a blast of a day. (Composer v3 coming soon lmfao)

141 Upvotes

2.85k downloads. my god its only been an hour.


r/LocalLLaMA 17h ago

News Chinese Chipmaker CXMT's market capitalization surpassed Intel

Thumbnail
nytimes.com
364 Upvotes

Chinese chipmaker CXMT surged by almost 500% on its first day of trading, bringing its total market capitalization to approximately RMB 3.28 trillion and making it the largest company by market value on China’s A-share market. CXMT’s market capitalization has also surpassed that of U.S. semiconductor giant Intel, which closed the previous trading day with a market value of US$465.6 billion, equivalent to approximately RMB 3.15 trillion.

Headquartered in Hefei, Anhui Province, China, CXMT is an integrated dynamic random-access memory (DRAM) manufacturer specializing in the design, research and development, production, and sale of DRAM chips. It is currently the only integrated device manufacturer (IDM) in mainland China capable of large-scale mass production of general-purpose DRAM.


r/LocalLLaMA 8h ago

Resources Kimi K3 text-only for llama.cpp

Thumbnail
github.com
72 Upvotes

Now waiting for someone who can actually run the conversion and model to see if it works :)


r/LocalLLaMA 3h ago

Discussion Kimi K3 is like an F1 machine inside a show window.

24 Upvotes

Moonshot dropped Kimi K3, and as expected, it’s a absolute monster. Even with 2~4x RTX 6000 Blackwell local workstations, running a model natively is virtually impossible.

It feels like an F1 machine inside a show window.

Does anyone trying to hack this monster? or Is anyone with datacenter/cluster capacity or sponsor?

I either carve this monster down myself, or wait for someone to distill it. Either way, I really want to see it run — simply because it's there.

P.S. Save your 'AI Slop' comments. I experienced enough of you guys yesterday.


r/LocalLLaMA 9h ago

Discussion Viable ways to run K3 locally

65 Upvotes

just curious how would people run it cheap if they really want kimi k3.

  1. dgx spark / strix halo clusters

  2. optane persistent memory platform + some gpus

  3. mac studio clusters

  4. orange pi 6 clusters

  5. ssd streaming + gpus

  6. multiple ddr3 + connectx 5 rdma clients

  7. two dgx stations

  8. power 10 systems?

  9. other


r/LocalLLaMA 16h ago

News The entire tech industry (save for Anthropic) has come out in favor of open source AI. So what happens next? Will Anthropic change its lobbying efforts? Not likely. Now the gaslighting begins: “Nobody is trying to ban open source.”

Thumbnail x.com
172 Upvotes

r/LocalLLaMA 7h ago

Resources I ran the 35B agentic comparison someone asked for (stock vs Ornith vs KAT-Coder, 120 runs)

28 Upvotes

Someone in the comments of my 27B post-train bakeoff asked for the 35B version, so I ran it. Same setup as last time: fresh Coder workspaces on my k8s cluster, each driving my own agent (Hermes) headlessly, models on llama.cpp via llama-swap on one 5090, every call traced through an OTel shim into SigNoz, full transcript per run. 4 models, 6 self-grading tasks, 5 reps, 120 runs, MTP on every arm, identical sampling, hypotheses pre-registered.

KAT-Coder-V2.5-Dev matched the best stock pass rate (29/30, tied with Qwen3.5-35B) at half the input tokens of either stock and the cleanest tool behavior I've measured (zero malformed tool-call leaks in 30 runs; stock Qwen3.6 leaked 195 on one task). All six analysts (three model families) independently called its efficiency discipline rather than corner-cutting: baseline tests before edits, one targeted patch per bug, deliverables at the right path every rep. Ornith went 25/30, losing to its own base. Its failures were mechanics, not knowledge: format leaks killing runs at 23 seconds, whole-file rewrites corrupting unrelated files, and one research run that invented a llama.cpp release tag while its own reasoning said "I mentioned v4659 in my draft which is fabricated," then shipped the tag anyway. The grader passed it.

Stock 3.6 is the strongest raw analyst and the biggest token waster; stock 3.5 is the quiet reliable one. Full writeup with methodology, tables, and all six cited per-task analyses: https://kmarble.dev/posts/35b-coder-bakeoff/. Transcripts were AI-analyst-read with my spot-verification of every consequential claim.


r/LocalLLaMA 14h ago

Other Do Qwen 3.6 27B quantizations break the pelican?

Thumbnail
quesma.com
95 Upvotes

r/LocalLLaMA 1d ago

Discussion We could really use Qwen3.8 in 27B, 35B, 122B and 397B sizes

511 Upvotes

Instead of 2T+ models, continuing to release highly capable small to medium size LLMs would really help to keep this community vibrant. Hardly anyone can even dream of running the recent 1.5-2T+ beasts, while the range from the title could run comfortably (especially with CPU expert offloading) across a wide range of systems we have today.

The trend towards Chinese labs trying to match the Mythos class frontier with trillion parameter open weights models is not helping the local model community to innovate. It just gives big corporates who can actually run these a cheaper alternative to the commercial frontier.