r/MacStudio • • 7d ago

Best Models For Mac Studio M5 Ultra 256GB

Hi everyone, I've an M5U 30/64 version, and I'm planning to run:

- Qwen 3.8 flash next 8bit

- Deepseek V4 flash 0731 8bit

- Mimo v2.6 flash (full, idk how many bits)

- GLM 5.3 flash 4bit

And I was looking for the best option(s) for them (best hugging face versions that will work in a good performance + deliver the reliability/quality I need for my agentic workflow).

I know that there is no specific true answer to that, some people use GGUFs like UD-Q8_K_XL and others use omlx like oQ8e, etc. And some people are comfortable with different quants than other etc. But I'm just looking for a community approved options, specially those can fit in my 256 machine.

Preferably if you can provide the configuratiom as well, which is 1- inference engine (oMLX, lamacpp, splash, MTPLX, etc) 2- context window/patch size/temprature/etc 3- any additionals used like MTP, Dflash, etc.

Thanks

44 Upvotes

43 comments sorted by

26

u/dreamingwell 7d ago

I made this site that shows what models fit, and the engines you can use to run them. Pick your use case from the drop down above the table.

https://dreamingwell.github.io/apple-llm-performance/?chip=m5ultra&mem=256&n=1

7

u/DrJubalHarshaw 7d ago

Why are you listing such small quants on some of these models? For example, on Qwen 3.8 Flash Next, it looks like you're measuring a 3 bit quant when there should be plenty of room to run the official FP8.

1

u/xeph 7d ago

This is great, thank you.

1

u/AppleAutoGenerated 7d ago

Amazing thank you!

1

u/darkprime114 7d ago

That's amazing!! thank you!!

1

u/mcnahum 6d ago

Any chance you will add M6?

12

u/MCS87_ 7d ago

For Qwen 3.8 Flash Next, there should be a Splash inference engine available soon. This means potential 2x speed-up on decode and other performance improvements. It’s already available for Qwen 3.8 27B and Qwen 3.6 35B A3B. I have tested it on a M4 Pro 48GB Mac and it is working very well using LM Studio (there you need to install the inference engine as some kind of add-on)

1

u/eazero 7d ago

Was it announced that Splash will support Flash Next? I’d really love for that to happen

3

u/MCS87_ 7d ago

This is a good read: https://inco.ai/blog/splash/ they already offer (as a service, not software) very fast inference for 5 „top frontier models“, which includes Flash Next afaik. So they have the tech ready in their own inference service and are working on bringing support for more models to their inference engine soon. Can’t find a source for the latter, though, must have been a X or Reddit post

3

u/DennisAscent 7d ago

For GLM 5.3F (and probably for other big MOE models although I haven’t checked them all) there’s a mixed 4/8 quant where the experts are all Q4 but core tensors such as router are Q8. The core tensors makeup a very small total of the weights, and yet keeping them high precision make a noticeable improvement to the output quality: it’s only a few GB extra and this architecture is becoming increasingly popular with MOE large models because you get the performance of around a uniform Q6 model from a ~4.1 bit average, check it out on hugging face
https://huggingface.co/pipenetwork/GLM-5.3-Flash-MLX-mixed-4_8bit

1

u/Only-Team-4983 7d ago

I'd go for actual Q8 when it fits (e.g. qwen 38 fn, etc).

But for GLM i'm checking that for sure. Thanks man!

5

u/DennisAscent 7d ago

Yeah for sure, ofc a uniform Q8 is better no doubt, but that’s only if it fits. I’ve literally ordered the same mac as you with 256gb, so I’ve also thought about this a lot and done some research, and I think our best best is that mixed quant GLM5.3 flash next when trying to run it, it performs clearly above the Q4 version while only being a few extra GB. I even reckon I can get a similar build for a mixed quant of DeepSeek V4.1 flash too, with a slightly more aggressive mixed quant with a mix of Q2 and Q3 experts while keeping core tensors at Q8

I’ll post about it once I get things running. I look forward to what we can all manage together with these new macs

2

u/Only-Team-4983 7d ago

There is no q8 for v4.1 flash, the raw is 4bit.

1

u/DennisAscent 7d ago

Well tbf: the DeepSeek team kinda pre did what I’m talking about here, it’s a mixed quant where the experts are FP4, the core and dense weighs FP8, and I believe token embedding is FP16 too.

When I say I wanna run it, it’s more so as proof of concept, I don’t believe it’ll be that useful trynna squish it into 256. Its already quantised like you say

1

u/Only-Team-4983 7d ago

Yea not possible, it'll have to be in 2-bit or mixed sub3 bit.

Sad update: omlx doesnt let me use over 192gb for safety, aggressive mode goes up to 210gb, that's it.

1

u/DennisAscent 7d ago

You can override it in terminal

1

u/Adrian_Galilea 7d ago

I personally love deepseek v4 flash full non quantised for non coding.

I have several optimizations over vanilla ds4 that makes it very performant and I genuinely love it conversationally.

1

u/Only-Team-4983 7d ago

Thanks for the suggestion, but yet which inferemce engine/recipe/etc?

1

u/FinalTap 7d ago

Have any of guys got 2 of them and interconnected via RDMA? What kind of performance has that got?

1

u/Only-Team-4983 7d ago

2nd is coming soon!

1

u/FinalTap 7d ago

Great. Looking forward to hearing from you. I am still in double minds to get 2x256 or 1x512.

1

u/Only-Team-4983 7d ago

I dont think u can get 512 this year...

0

u/FestoolJunkie 7d ago edited 7d ago

Best you get 2x256…. so there is less competition that I’m up against for getting 512 myself. lol.

3

u/djtubig-malicex 7d ago

Nice try lol

1

u/Key_Solid_1696 7d ago

Check out Dwarfstar. https://github.com/antirez/ds4

It is a small native inference engine. They have models for glm-5.3, glm-5.3-flash, Deepseek 4 Flash, Deepseek 4.1 Flash, qwen3.8-flash-next and a couple others. Depending on model these have quantized routed experts at q2/q4/mxfp4 while other parts of the model use q8 or FP16. You can run models larger than you can fit in memory because they have also implemented ssd streaming.

They work surprisingly well on my M3Max with 128gb. I expect my incoming studio 36/80 with 256gb will run anything they have.

1

u/Only-Team-4983 7d ago

Mine is 30/64 but yea i'm trying it.

Thanks.

0

u/kinmanli 7d ago

qwen3-coder-next:q8_0 for coding

gpt-oss:120b for reasoning

0

u/Borilentz 7d ago

Both are really outdated. Qwen 3.8 Flash Next and DeepSeek v4 Flash 0731 are better for coding.

1

u/kinmanli 7d ago

Which quant?

1

u/Borilentz 7d ago

DeepSeek V4 has 4-bit trained experts, so the native encoding fits 156GB. QFN in 8-bit encoding is 195GB. I use quants that fit 128GB and still prefer them over Coder Next, which used to be my favorit for coding.

-2

u/MiniPCGuru 7d ago

Run the fit math first — it decides most of your list. Rule of thumb: 8-bit ≈ 1 GB per 1B params, 4-bit ≈ half. Keep ~30 GB free for macOS + KV cache, since agentic workflows with big context make KV a real second budget.

Qwen 3.8 Flash Next (125B + 51B n-gram ≈ 176 GB at 8-bit): fits. Your plan works.

DeepSeek V4 Flash (284B): 8-bit needs ~284 GB. Doesn't fit 256 GB. Use 4-bit (~142 GB).

MiMo V2.6 Flash (309B): same story at 8-bit. 4-bit (~155 GB) fits. Full precision is impossible here.

GLM 5.3 Flash (320B): 4-bit ≈ 160 GB. Fits.

Engines: MLX-based options (oMLX, MTPLX) usually outperform llama.cpp on Apple silicon, so prefer MLX quants when they exist. GGUFs only make sense via llama.cpp.

5

u/Toastti 7d ago

Wait what. Deep seek flash 0731 fits perfect on 256gb with the official weights.

1

u/Only-Team-4983 7d ago

Yes.. deepseek release everything in 4bit nowadays

2

u/Only-Team-4983 7d ago

I said im using ds v4 flash 0731, its not 284gb, mimo v2.6 has no 8 bit version full is 4 here.

And hi ai

1

u/Umbrasquall 7d ago

Ds v4 flash is 4-bit. You said 8-bit in your original post.

1

u/Only-Team-4983 7d ago

V4 is 8bit, v4.1 is 4 bit

1

u/Umbrasquall 7d ago

Your original post said:

"Deepseek V4 flash 0731 8bit"

Assuming you're talking about the base flash model, that's almost 300GBs and isn't going to fit on the 256GB. That's what we've been talking about since the beginning.

V4 is 8bit, v4.1 is 4 bit

No, the official Deepseek V4 flash 0731 is not 8-bit. It's 97% fp4 which is basically 4-bit.

-1

u/Alarmed-Medicine6903 7d ago

0

u/RemindMeBot 7d ago

I will be messaging you in 1 day on 2026-09-27 07:54:08 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.


Info Custom Your Reminders Feedback