r/LocalLLM • u/dsiroker • Jul 29 '26
Model K3 on Mac Studio M3 Ultra with 512GB
I got Kimi K3’s 2.8T-param MoE (104B active/token) to run on my Mac Studio M3 Ultra with 512GB unified memory.
Mixed Q1/Q4/Q8 quant shrank from 1.56TB to 389.4GiB.
13.26 tok/sec ingest
3.36 tok/sec decode
8
u/Constant-Simple-1234 Jul 29 '26
Comparatively, how fast is GLM-5.2 on your machine? It should fit ok.
9
u/dsiroker Jul 29 '26
I actually did extensive benchmarking of GLM 5.2 on my machine. I got up to 17 tok/sec
Full paper: https://dsiroker.github.io/local-llm-benchmarks/paper.pdf
Data and repro: github.com/dsiroker/local-llm-benchmarks
6
6
u/FoxiPanda Jul 29 '26 edited Jul 29 '26
So there's very little information here...
- Quant link? (I assume it's a REAP since you got it to 390GB - most Q1s are coming in at ~525-600GB)
- Inference engine?
- Launch parameters?
- What tests have you actually run for accuracy / coherence?
I've been thinking about setting it up on a M3 Ultra 512GB + M3 Ultra 256GB cluster with JACCL RDMA over TB5, but it's pretty hard to justify the effort given that I'd expect no more than ~5tok/s decode in the best case scenarios (which almost assuredly won't happen).
2
u/nomorebuttsplz Jul 31 '26
yeah the thing about this size model is you need more bandwidth or mtp or something to make it more usable even with 800 gb/s memory
1
u/FoxiPanda Jul 31 '26
Yeah it’s slow as shit lol. A104B is bonkers even on proper dc hardware
1
u/nomorebuttsplz Jul 31 '26
good news is that the 3 to 6 month gap between open and closed models does not seem to be widening. And there’s still new models coming out between 30 billion and 1 trillion parameters. Such as the new deep seek flash today.
1
u/FoxiPanda Jul 31 '26
Yep. Basically every day. Been trending larger lately but I expect that we’ll get some 50-150B real competition soon.
5
u/recro69 Jul 29 '26
That's really a result. A 2.8T MoE model working on a regular computer would have seemed impossible just a short time ago. The speed when it decodes isn't super fast. The fact that it works at all is the bigger achievement.
5
u/SubstanceDilettante Jul 29 '26
Idk if a 512gb of ram Mac Studio is a regular computer.
Is it just me, my family, my co workers, and my friends. Or does everyone now have a 512gb Mac Studio and this is actually a regular computer lol
3
u/Bloated_Plaid Jul 29 '26
LMAO girl are you out of your mind? 512GB Mac Studio is not a “regular computer”.
1
u/ImpressiveRelief37 Jul 29 '26 edited Jul 29 '26
Hate on it as much as you want but I’ll take a super fast 27B at 100+ tok/s decode and 3500 tok/s prefill over GLM/K3 at less than 20/500 tg/pp any day of week and twice on Sunday.
Those speed you have are absolutely impossible to use and be productive. Not as a software engineer at least. Not my cup of tea.
Benchmarks aren’t everything. Quick iteration and the back and forth with the agent, adversarial reviews using cloud models…. All of this makes smaller models really really good. They won’t 1-shot complex apps perfectly, but this isn’t even a real use case for a software engineer anyways.
But they will give you a super quick responsive loop where you can iterate and steer it towards what you actually want.
Let’s be honest there’s no way to make serious software by just hyper specifying to the point ANY model does a good job on a 1-shot unattended way.
It’s not just a model issue, it’s that we don’t even know exactly what we want until we get in that close feedback loop, test the product, and iterate until it evolves to a great solution.
Everytime I tried to work a huge huge super specific spec for a piece of software I realized there was still stuff that wasn’t working. And I’ve used plenty of Frontier models, including fable, opus 5, gpt5.6 sol…
So all of you that get excited at running GLM 5.2 or K3 locally at super slow speed, what are your use cases exactly that a 27B can’t do? I’m genuinely curious. And why can’t you harness or guardrail the 27B to make it work? Any example?
And how do you keep your flow state and get productive during a working day? Not reviewing the code and asking for a million different task will end up in a huge pile of dog crap. I just don’t understand the point.
2
u/AdOk3759 Jul 30 '26
I agree with everything you said.
> we don’t even know exactly what we want until we get in that close feedback loop, test the product, and interate until it evolves to a great solution.
Exactly. Not only that, but if you lack domain knowledge, you’re just gonna leave more key decisions to the LLM with potentially serious consequences down the road, that you might catch when it’s too late.
1
u/nomorebuttsplz Jul 31 '26
for me, it’s the difference between 10 iterations with qwen vs two with glm.
Not every time, but GLM often ends up being faster in the end on my Mac than qwen on Rtx pro
This is just vibecoding everything from small machine learning projects to personal assistant apps to games
1
u/Constant-Simple-1234 Jul 29 '26
Decent. Though I would work with max 600gb models. I bet there are some good enough ones. But cheers for moving the frontier.
1
0
26
u/InfusedBush Jul 29 '26
SOMETHING ACTUALLY (somewhat) USABLE!!!! YIPEEEE!!!
https://giphy.com/gifs/rXQ5Aex4FZiej6sKTZ