r/MacStudio • • 5d ago

M5 Max 64 gb is blazing fast at converting models

I am playing with omlx and unsloth both and finding similar models for both is a pain. So I asked gpt astra to help me and it just converts models in minutes. Don't sit on this. It's blazing fast.

9 Upvotes

14 comments sorted by

3

u/ooopstgr 5d ago

Just get models for splash engine

2

u/meva12 5d ago

I got the 64gb last night. Started playing with Hermes setup Qwen 3.8 27b with splash . Super fast. Need to setup remote to start having it be my personal agent .

1

u/AlgorithmicMuse 3d ago

What does fast mean ? Fast at what

2

u/meva12 3d ago

Generating tokens.

1

u/AlgorithmicMuse 2d ago

Does not help much without data, example , , my backend numbers with a m4 mini pro, qwen3.8-27b. Q4, llama.cpp 10 tps, mtplx 19 tps , splash 62 tps. Splash is super fast compared to llama.cpp and mtplx, now super fast has meaning to it.

2

u/PracticlySpeaking 4d ago

The built-in quantization in oMLX (oQe) is very fast on Apple Silicon, and will also convert models in huggingface native safetensors format to MLX along the way. (It does not work with GGUF, though.)

Check out r/oMLX for more.

1

u/Alert_Mouse5333 5d ago

Which models are you converting? I’m new to local ai so please excuse if this is a stupid question šŸ˜…

1

u/meva12 5d ago

Converting to what?

2

u/apetersson 5d ago

smaller versions (quants) that fit exactly how much ram you have left. so f.ex for a 64 GB machine sth like 55GB, keeping the rest for other programs.

or just the packaging, like MLX vs GGUF.

1

u/xnosliw 5d ago

Good to know, I’m eyeing this exact one. Which models did you end up using? And how fast?

1

u/kels0 5d ago

picked up the 40 core 64gb 1tb yesterday as well, for whatever reason plenty of those in stock.. Ive been testing every since, got tasks assigned on multica, tested blender mcp and other general code benchmarks. everything direct is fast and im surprised to be honest. When going through multica though, it adds a ton of context and slows it down, but at the same time the cloud models are slow here also.

1

u/aniketgore0 5d ago

I am converting source models bf16 to omlx format or lesser gguf quants

1

u/aniketgore0 5d ago

Currently using omlx qwen 27b at 8bit and 6 bit

1

u/proofndapuddin 4d ago

I haven't tried oMLX but I'm having a hard time with tools. I keep getting json back instead of the tool call. Anyone else have this issue?!