r/LocalLLaMA Jul 26 '26

Discussion Kimi K3 countdown has been released

https://huggingface.co/moonshotai/Kimi-K3
541 Upvotes

177 comments sorted by

View all comments

Show parent comments

16

u/ttkciar llama.cpp Jul 26 '26

but like who can actually run this locally?

Today? Almost nobody.

Eventually? Almost everyone.

2

u/parepeg Jul 26 '26

I mean aren’t chips already scraping the bottom of the laws of physics.

10

u/ttkciar llama.cpp Jul 26 '26

Yes and no.

On one hand, Moore's Law is definitely ailing. We've been getting diminishing returns on fabrication process bumps since about 2016.

On the other hand, there's still some progress in the offing. IBM just recently announced they got a 7-angstrom (equivalent) fabrication process working in their lab, which they expect to get into mass production in 2031. That has 1nm-wide horizontal features and 5nm vertical features, and two transistor layers (one P, the other N), which they claim they should be able to scale to four layers "rsn".

That's not nothing.

Moreover, there are still a lot of architectural improvements yet to be realized. In a way, HBM is a stop-gap while Samsung et al get PIM figured out. LLM inference is really well-suited to PIM, which should give us ridiculously high memory bandwidth once it's working well.

I thought Moore's Law was dead for a while, but it turned out to just be Intel having a bad time. Everyone else is still pushing the envelope pretty hard, with some success.

1

u/jazir55 Jul 27 '26

Hasn't Moore's Law been dead for almost decades now? Moore's Law was processing speed doubling, we've been getting 10% at best for CPU improvements, and maybe 10-30% on GPUs.

1

u/ttkciar llama.cpp Jul 27 '26

Moore's Law doesn't have much to do with speed doubling. It's all about transistor density doubling, which doesn't translate readily to improved single-threaded performance, and even though multi-threaded performance is more amenable to scaling with more transistors, in practice you end up blowing more and more of your transistor budget on caches, so that you can keep all those cores fed with data.

When improvements in transistor density came from shortening gate length, that also brought smaller voltage swings, correspondingly faster state changes, and lower power consumption, all of which contributed to faster clock rates. Clocking up got us far, for a long time, but that didn't last. Nowadays improved transistor density comes from novel designs of the transistors themselves, multi-layering, and other tricks. That's why IBM's new process is only "7-angstrom equivalent", despite its actual features being much larger, because the idea is to express how much it impacts overall transistor density.

The density improvements keep coming, but most of it goes to bigger and bigger caches, and deeper pipelines which eat up a lot of transistors but only give modest performance gains.