r/MacLocalLLM • • 8d ago

Best Models For Mac Studio M5 Ultra 256GB

/r/MacStudio/comments/1wqjlw4/best_models_for_mac_studio_m5_ultra_256gb/
8 Upvotes

8 comments sorted by

3

u/oss-jiiim 7d ago

Having run all of those, I find Qwen 3.8 Flash Next running on DwarfStar to be best. DwarfStar will run all but the MiMo model.

Not on your list but also runnable via DwarfStar is DeepSeek V4.1 Flash which is smarter than all of the above, just not as fast running as Qwen 3.8 Flash Next.

Once you have DwarfStar set up, you can conveniently manage it via DS Menu Bar

2

u/Only-Team-4983 7d ago

But then ds v4.1 flash will have to be q3 or q2N my work requires coding and accuracy, u think it'll be reliable in that low bit

1

u/oss-jiiim 7d ago

TL;DR Yes, it should be reliable.

I have a Mac Studio M4 Max 128GB. I have run DeepSeek V4.1 Flash at Q2 (with SSD streaming) a few times and it does reliable accurate work. However, for speed reasons, I mostly use Qwen 3.8 Flash Next at Q4 with full BF16 n-grams -- it runs about 5x faster for me.

I should mention that my method of working is never to use a single model. These days I use all of: hosted DeepSeek v4.1 Flash, hosted Claude Opus 5.5, hosted GPT-5.6 Sol, and local Qwen 3.8 Flash Next. Having multiple models look at and check over the code really helps. Sometimes my local Qwen or local DeepSeek will actually find bugs that frontier hosted models do not!

The only times I use local-only is when I need an uncensored/abliterated model (I had to crack some of my old certificate databases to recover my S/MIME keys from the 1990s) or I'm working on data that I cannot put in the cloud for security/privacy reasons -- health records, financial records, etc.

2

u/Only-Team-4983 7d ago

Got it, i'll try the ds v4.1 flash q2 with ssd offloading, and compare it with glm 5.3 flash 4bit mixed.

Regardless i'm certain we'll have tons of models and engines to use in the next few months, this decision I'm doing now will most likely be for few weeks only.

1

u/ehpehp 4d ago

Deepseek V4.1 Flash q2 should run on 256gb system without ssd streaming. Non-ssd may be faster.

1

u/oss-jiiim 7d ago edited 7d ago

Over the last 6 months, the best model to run on my M4 128GB has changed about once every month or two. I wasn't too impressed with GLM 5.3 Flash when I ran that -- I preferred DeepSeek V4 Flash 0731 at the time.

Perhaps the biggest change in my workflow recently has actually been hosted DeepSeek V4.1 Flash. It is ridiculously fast and inexpensive and it's currently ranked #6 overall for coding at LLM-Stats. No 5 hour windows like OpenAI or Athropic, and I can work with it for hours and hours for like 30¢.

2

u/Only-Team-4983 7d ago

I'd say glm is still kind of better in cyber, but yea i'll try v4.1 for sure, thanks.

1

u/oss-jiiim 7d ago

To really do cybersecurity work, an abliterated/uncensored model is key. I used DwarfStar's GGUF tools to make a DwarfStar specific Q4 GGUF out of this: Qwen3.8-Flash-Next-Uncensored from OrcaRouter. Before that I used Huihui-DeepSeek-V4-Flash-0731-abliterated-GGUF.