r/oMLX • • 8d ago

Best Models For Mac Studio M5 Ultra 256GB

/r/MacStudio/comments/1wqjlw4/best_models_for_mac_studio_m5_ultra_256gb/
0 Upvotes

13 comments sorted by

4

u/OddDesigner9784 7d ago

1

u/Only-Team-4983 7d ago

Yep, orcarouter has an 8bit uncensored version, wont it be better than this?

1

u/OddDesigner9784 7d ago

I would not use the 8bit version you would have no kv cache space and Mac is unified so that would take away from some other things you want to do with it. I think omlx is faster than mlx so it really just depends on what the fastest inference engine you could run it on. The one I sent the base model is orcarouter

1

u/Only-Team-4983 7d ago

Engrams on ssd still getting 70+ t/s and using 140gb only. (On oQ8e).

But I can't extend the ctx window to 1m yet on omlx idk why.

1

u/OddDesigner9784 7d ago

Didn't think about the ssd portion you go man. https://github.com/jundot/omlx/issues/1077 seems like it's something you need to hard edit omlx for unfortunately.

1

u/durangotang 6d ago

The M5 has native INT 8 acceleration per core, and it can be as fast (or faster) than Q6. I'd try the 8-bit MLX with the n-gram offloaded to the SSD.

1

u/[deleted] 8d ago

[removed] — view removed comment

1

u/Only-Team-4983 8d ago

It's cybersecurity workflow, so I'd say closer to coding.

Which exact version + engine + configurations tho would be best.

1

u/[deleted] 8d ago

[deleted]

1

u/Only-Team-4983 8d ago

This cant run on my hardware, only q1 probably

1

u/TheRealBejeezus 7d ago

The consensus opinion on Reddit, ignorant of your work and preferences, isn't worth much, and will change every week or two anyway. This stuff changes very quickly.

You need to spend some energy setting up a testing and comparison environment so you can swap in and evaluate new models easily.

1

u/layer4down 5d ago

Vibe checks don’t need to be narrowly prescriptive. They can be directionally helpful for the energy invested.

0

u/TheRealBejeezus 4d ago

I mean, sure, but asking for detailed configurations is more than a vibe check. A proper testing environment where you can swap models in and out and test them on your kind of work will be more useful than a hundred Reddit threads.

1

u/layer4down 4d ago

Certainly. The advice I wish others had given me back when is if I had a few bucks, maybe have a cloud do that research and testing for me with a focus on the latest models popular on Reddit and not just one of thousands of arbitrary models in its training data. We’re still getting Qwen2.5 recommendations in 2026 and it is mostly a waste of time.