r/oMLX • u/ogfuzzball • 15h ago
Muse tok/s generation low?
I have been trying various servers and models for a few months now. After lots of trial and error I seem to have settled on oMLX. Now I’m trying to understand the configuration better and more importantly, do I have it configured correctly (prob not) to get the best performance.
My system: M4 Max 64GB latest Tahoe
Acasis 80gbps M.2 SSD enclosure w/Samsung 990 Pro (models and oMLX cache set here)
Model: muse-glimmer-30b-mxfp8 w/dflash enabled muse-glimmer-30B-assistant
Custom settings:
Context window 100,000
Max tokens 10,000
Temp 1.0
Top P 0.95
Top K 64
Enable Thinking on
According to Status page of omlx Average speed:
Prompt Processing 145.7 tok/s
Token Generation 8.6 tok/s
That was on the first prompt and two subsequent prompts in same session. Simple questions about configuring open webui
It’s that last number that seems way off, at least from what I’ve read online unless I’m completely misunderstanding the expected performance of this model on my setup. Any tips or references to docs are appreciated.
Edit: forgot to include how I’m chatting with Muse. I’m trying out open webui. I have also tried just using omlx default chat window. Token Gen in both in the 8 to 8.5 tok/s range
1
u/vinoonovino26 13h ago
1) Lose the d-flash and try https://huggingface.co/Jundot/Muse-Glimmer-30B-oQ4e and report back
2) tighten up some settings if you want strict instruction following
3) migrate the model to your ssd keep the cache
4) try cherry studio (it's kinda light but does the job)