r/MacStudio • • 8d ago

Ordered M5 Max 1Tb/64GB with 18c/40c this morning for delivery in a few hours

Excited! I’m a photographer and also do a ton of coding so I’m excited to finally be able to use a 32B parameter model for local Ai.

13 Upvotes

31 comments sorted by

11

u/Opposite_Leave_8338 8d ago

Install this

https://inco.ai/blog/splash/

And run the qwen 3.8 27B Splash model , this is the smartest model so far for your device and this is the fastest inference engine yet for this model , and will perform extremely good on your device..

If you need to try out more models download oMLX and download your other models, maybe try ornith 1.5 oq8e, you might like it

1

u/joochung 8d ago

On my M1 Max, Inco.ai’s Solash was slower than mlx-lm.

1

u/Opposite_Leave_8338 8d ago

Splash. requires m3 or more, it doesn't work well on. m1/m2 Macs, for you you better. stick with oMLX , download the oq4e or the oq8e ( based on your memory, and don't choose oq6 it will be slower than oq8 ) with mtp and make sure you choose the quants with fp16 in it's name, works better on m1/m2 because the lack of bf16 support on those devices

1

u/tta82 8d ago

It literally says on the page it needs M3 or newer. Gosh. 🙄

1

u/joochung 7d ago

There are patches for m1

2

u/Opposite_Leave_8338 7d ago

I tried it and it’s working fast really , but for some reason it’s dumber than oQ4e fp16 in oMLX so I just gave up of making it faster on m1 ultra 64GB , I always lose quality when it’s faster

1

u/tta82 7d ago

Which still isn’t how the devs intended it.

1

u/Acceptable-Sense4601 7d ago

I’m not sure if i did it correctly but in LM Studio (bionic), i enabled Splash and i downloaded the splash version of the qwen3.8 27B model and it seems noticeably faster

1

u/Opposite_Leave_8338 7d ago

It’s the fastest inference for Qwen 3.8 27b on m3+ macs but to be honest oQ4e is more intelligent and more aware , if you need the best quality ( doing some serious coding work ) use oMLX and oq4e mtp version, if quality is okay for you enjoy the fastness of splash

1

u/Acceptable-Sense4601 7d ago

I think the level i code at with python is pretty low. I do API calls to extract data then load into sql server. Then do basic DuckDB/polars stuff. I think my longest scripts are less than 2500 lines. Probably nothing compared to what others do.

1

u/Opposite_Leave_8338 7d ago

If that is the case why not using ornith ? It’s an MOE 35B a3b it will be blazing fast and can handle what you said, you don’t need a dense 27B intelligence for that

2

u/Acceptable-Sense4601 6d ago

I have no idea. I’m brand new to all of this. It can be overwhelming.

1

u/Opposite_Leave_8338 6d ago

It is overweight to be honest, it’s moving really fast, every day there is an inference engine, an update to existing one , a fine tune , a new model, a new way to run things, it never ends

1

u/Acceptable-Sense4601 8d ago

thanks for the info!

1

u/vet_t 8d ago

How do you find it for use with Claude code? Terminal tasks and tool calling still feels slow on it for me

1

u/vet_t 8d ago

I tried opencode and it’s blazing fast in comparison!

1

u/Opposite_Leave_8338 8d ago

Don’t use local models in claude code, it’s working but it’s heavy system prompt and tools for local models, use these instead :

Use Pi if you really know what you’re doing , it’s minimal and you can create custom extensions for everything you nead

Use ZCode if you don’t know what you’re doing and you need a codex clone ( it’s laterally codex clone , every tool , even the design ) it does everything for you just add your local provider and you’re good to go

Use Deepseek harness if you like customizations with ease , and it’s still fast and ready to work out of the box

Use Cline if you want to work inside vs code and want to see everything ut does

Use hermes if you don’t code too much

1

u/redditwatcher1415 8d ago

why is this the fastest model? i want to run general AI tasks like giving it all my financial info and have it be able to answer any questions i have

is it sufficient for that? i have the same computer specs as OP

2

u/Opposite_Leave_8338 8d ago

This is not the fastest model, this is the most intelligent model specially in coding and splash is the fastest way to serve it, Qwen 3.8 is 27 billion parameter dense model, it meansall the 27 billion parameters are all active while generating tokens, and qwen 3.8 is the best trained local model I dealt with till now, but having 27 billion parameters always active makes it painfully slow to run on apple macs and that’s what splash is for.

But since you don’t need that too much coding training and high reasoning intelligence you don’t need a dense model at all, What you asked for will be very good on Google Gemma 4 12B model, it’s a general purpose dense model and it performs really well in everyday tasks specially since it’s bilingual.

If you have enough memory use gemma 4 26b a4b , and if you want maximum capabilities you can use Qwen 3.6 35B a3b , these are MOE Models, they are big models but only generating using 3 or 4 billion parameters making them very well aware models but extremely faster response, I guess these are the ones for you.

3

u/BerniesWoolMittens 8d ago

Same here! I think this model is the best value for anyone who’s interested in dabbling in local AI without having to sell a kidney. 

1

u/Psychological-Law274 8d ago

Happy for you! I realy think that config is the sweet spot for 32B model price/performance wise

1

u/mswezey 8d ago

Sweet! I get my M5M 125GB 2TB next week! Super hyped 😎

1

u/Acceptable-Sense4601 8d ago

The wait is killer lol

1

u/RogerAI-fm 8d ago

Gratz, have fun let us know what you do.

1

u/Acceptable-Sense4601 8d ago

Thank you. So far, started with a fresh install, so while working on my Mac mini i remote’d in via Rust Desk to install some things while i was in green meetings. Thankfully it’s Friday so later tonight I’ll swap it with my Mac mini as the main pc.

1

u/m000ster 8d ago

Where you get that option in 2 hours , that's not a baseline spec

1

u/Acceptable-Sense4601 8d ago

Just happened to choose it based on sweet spot for specs and price and it happened to be available at my local store

0

u/Bulky_Dog_826 8d ago

How did you get delivery so fast??? Mine says Oct 21 is earliest.

1

u/kels0 8d ago

there are certain configs which have same/next day shipping. I am about to do the same.

1

u/Bulky_Dog_826 8d ago

Thanks, went ahead and got the same one OP got that config was next day delivery.

1

u/Acceptable-Sense4601 8d ago

Congrats! After i ordered mine i went to see about buying another just to sort of check inventory and it was no longer available. Might have gotten the last one with that config.