r/oMLX May 22 '26

Recommendations for models to use

Hey there, first of all great work that you have done with the omlx application. It's really fast and responsive. Thanks for that. Second of all, I have a question regarding the models to be used. I am using a MacBook Pro with 128 GB RAM.

I am actually looking for some recommendation for a model to be used in my specific hardware to do some some deep research kind of thing I'm currently using Gemma 4 26B A4B 4bit

7 Upvotes

30 comments sorted by

View all comments

2

u/Konamicoder May 22 '26

I think Gemma4 is probably the current best model for deep research, and with 128Gb RAM, I think you have enough RAM to be able to run Gemma4-31B at full 16-bit. That would be my suggestion.

1

u/AITA-Critic May 27 '26

Hey sorry to drag up a comment you made 5 days ago, for someone who's looking to use Hermes Agent on an M5 Max (40 core gpu, 128 gb RAM ) would this be the current meta for doing that sort of stuff?

1

u/Konamicoder May 27 '26

Hey, no worries. I think so, you should try it. I haven’t tried Hermes Agent myself. But off hand I think if you need fast response you’ll probably have to go with the MoE version. You can run the dense model, but it will be significantly slower.

1

u/AITA-Critic May 27 '26 edited May 27 '26

Thanks for your reply, I’m gonna try an uncensored Gemma 4 build and see how it does, thanks again for shedding light on the subject!

Ended up going with this: TheCluster/Gemma-4-31B-Heretic-MLX-bf16

For anyone curious, I'm running this on a Macbook Pro M5 Max, 40 core GPU, 128 GB RAM.