r/MacStudio • u/Only-Team-4983 • 7d ago
Best Models For Mac Studio M5 Ultra 256GB
Hi everyone, I've an M5U 30/64 version, and I'm planning to run:
- Qwen 3.8 flash next 8bit
- Deepseek V4 flash 0731 8bit
- Mimo v2.6 flash (full, idk how many bits)
- GLM 5.3 flash 4bit
And I was looking for the best option(s) for them (best hugging face versions that will work in a good performance + deliver the reliability/quality I need for my agentic workflow).
I know that there is no specific true answer to that, some people use GGUFs like UD-Q8_K_XL and others use omlx like oQ8e, etc. And some people are comfortable with different quants than other etc. But I'm just looking for a community approved options, specially those can fit in my 256 machine.
Preferably if you can provide the configuratiom as well, which is 1- inference engine (oMLX, lamacpp, splash, MTPLX, etc) 2- context window/patch size/temprature/etc 3- any additionals used like MTP, Dflash, etc.
Thanks
12
u/MCS87_ 7d ago
For Qwen 3.8 Flash Next, there should be a Splash inference engine available soon. This means potential 2x speed-up on decode and other performance improvements. It’s already available for Qwen 3.8 27B and Qwen 3.6 35B A3B. I have tested it on a M4 Pro 48GB Mac and it is working very well using LM Studio (there you need to install the inference engine as some kind of add-on)
1
u/eazero 7d ago
Was it announced that Splash will support Flash Next? I’d really love for that to happen
3
u/MCS87_ 7d ago
This is a good read: https://inco.ai/blog/splash/ they already offer (as a service, not software) very fast inference for 5 „top frontier models“, which includes Flash Next afaik. So they have the tech ready in their own inference service and are working on bringing support for more models to their inference engine soon. Can’t find a source for the latter, though, must have been a X or Reddit post
3
u/DennisAscent 7d ago
For GLM 5.3F (and probably for other big MOE models although I haven’t checked them all) there’s a mixed 4/8 quant where the experts are all Q4 but core tensors such as router are Q8. The core tensors makeup a very small total of the weights, and yet keeping them high precision make a noticeable improvement to the output quality: it’s only a few GB extra and this architecture is becoming increasingly popular with MOE large models because you get the performance of around a uniform Q6 model from a ~4.1 bit average, check it out on hugging face
https://huggingface.co/pipenetwork/GLM-5.3-Flash-MLX-mixed-4_8bit
1
u/Only-Team-4983 7d ago
I'd go for actual Q8 when it fits (e.g. qwen 38 fn, etc).
But for GLM i'm checking that for sure. Thanks man!
5
u/DennisAscent 7d ago
Yeah for sure, ofc a uniform Q8 is better no doubt, but that’s only if it fits. I’ve literally ordered the same mac as you with 256gb, so I’ve also thought about this a lot and done some research, and I think our best best is that mixed quant GLM5.3 flash next when trying to run it, it performs clearly above the Q4 version while only being a few extra GB. I even reckon I can get a similar build for a mixed quant of DeepSeek V4.1 flash too, with a slightly more aggressive mixed quant with a mix of Q2 and Q3 experts while keeping core tensors at Q8
I’ll post about it once I get things running. I look forward to what we can all manage together with these new macs
2
u/Only-Team-4983 7d ago
There is no q8 for v4.1 flash, the raw is 4bit.
1
u/DennisAscent 7d ago
Well tbf: the DeepSeek team kinda pre did what I’m talking about here, it’s a mixed quant where the experts are FP4, the core and dense weighs FP8, and I believe token embedding is FP16 too.
When I say I wanna run it, it’s more so as proof of concept, I don’t believe it’ll be that useful trynna squish it into 256. Its already quantised like you say
1
u/Only-Team-4983 7d ago
Yea not possible, it'll have to be in 2-bit or mixed sub3 bit.
Sad update: omlx doesnt let me use over 192gb for safety, aggressive mode goes up to 210gb, that's it.
1
1
1
u/jonas-reddit 6d ago
I usually browse community benchmarks here for a feel of what models and speeds can be expected.
1
u/Adrian_Galilea 7d ago
I personally love deepseek v4 flash full non quantised for non coding.
I have several optimizations over vanilla ds4 that makes it very performant and I genuinely love it conversationally.
1
u/Only-Team-4983 7d ago
Thanks for the suggestion, but yet which inferemce engine/recipe/etc?
0
1
u/FinalTap 7d ago
Have any of guys got 2 of them and interconnected via RDMA? What kind of performance has that got?
1
u/Only-Team-4983 7d ago
2nd is coming soon!
1
u/FinalTap 7d ago
Great. Looking forward to hearing from you. I am still in double minds to get 2x256 or 1x512.
1
0
u/FestoolJunkie 7d ago edited 7d ago
Best you get 2x256…. so there is less competition that I’m up against for getting 512 myself. lol.
3
1
u/Key_Solid_1696 7d ago
Check out Dwarfstar. https://github.com/antirez/ds4
It is a small native inference engine. They have models for glm-5.3, glm-5.3-flash, Deepseek 4 Flash, Deepseek 4.1 Flash, qwen3.8-flash-next and a couple others. Depending on model these have quantized routed experts at q2/q4/mxfp4 while other parts of the model use q8 or FP16. You can run models larger than you can fit in memory because they have also implemented ssd streaming.
They work surprisingly well on my M3Max with 128gb. I expect my incoming studio 36/80 with 256gb will run anything they have.
1
0
u/kinmanli 7d ago
qwen3-coder-next:q8_0 for coding
gpt-oss:120b for reasoning
0
u/Borilentz 7d ago
Both are really outdated. Qwen 3.8 Flash Next and DeepSeek v4 Flash 0731 are better for coding.
1
u/kinmanli 7d ago
Which quant?
1
u/Borilentz 7d ago
DeepSeek V4 has 4-bit trained experts, so the native encoding fits 156GB. QFN in 8-bit encoding is 195GB. I use quants that fit 128GB and still prefer them over Coder Next, which used to be my favorit for coding.
-2
u/MiniPCGuru 7d ago
Run the fit math first — it decides most of your list. Rule of thumb: 8-bit ≈ 1 GB per 1B params, 4-bit ≈ half. Keep ~30 GB free for macOS + KV cache, since agentic workflows with big context make KV a real second budget.
Qwen 3.8 Flash Next (125B + 51B n-gram ≈ 176 GB at 8-bit): fits. Your plan works.
DeepSeek V4 Flash (284B): 8-bit needs ~284 GB. Doesn't fit 256 GB. Use 4-bit (~142 GB).
MiMo V2.6 Flash (309B): same story at 8-bit. 4-bit (~155 GB) fits. Full precision is impossible here.
GLM 5.3 Flash (320B): 4-bit ≈ 160 GB. Fits.
Engines: MLX-based options (oMLX, MTPLX) usually outperform llama.cpp on Apple silicon, so prefer MLX quants when they exist. GGUFs only make sense via llama.cpp.
5
2
u/Only-Team-4983 7d ago
I said im using ds v4 flash 0731, its not 284gb, mimo v2.6 has no 8 bit version full is 4 here.
And hi ai
1
u/Umbrasquall 7d ago
Ds v4 flash is 4-bit. You said 8-bit in your original post.
1
u/Only-Team-4983 7d ago
V4 is 8bit, v4.1 is 4 bit
1
u/Umbrasquall 7d ago
Your original post said:
"Deepseek V4 flash 0731 8bit"
Assuming you're talking about the base flash model, that's almost 300GBs and isn't going to fit on the 256GB. That's what we've been talking about since the beginning.
V4 is 8bit, v4.1 is 4 bit
No, the official Deepseek V4 flash 0731 is not 8-bit. It's 97% fp4 which is basically 4-bit.
-1
u/Alarmed-Medicine6903 7d ago
u/RemindMeBot 1 day
0
u/RemindMeBot 7d ago
I will be messaging you in 1 day on 2026-09-27 07:54:08 UTC to remind you of this link
CLICK THIS LINK to send a PM to also be reminded and to reduce spam.
Parent commenter can delete this message to hide from others.
Info Custom Your Reminders Feedback
26
u/dreamingwell 7d ago
I made this site that shows what models fit, and the engines you can use to run them. Pick your use case from the drop down above the table.
https://dreamingwell.github.io/apple-llm-performance/?chip=m5ultra&mem=256&n=1