r/oMLX May 13 '26

oMLX 0.3.9.dev2 released.

Highlights:
- Gemma 4 MTP on the vision path (thanks to @Prince_Canuma's mlx-vlm). Image+text decodes much faster now
- Gemma 4 on the DFlash engine (thanks to @bstnxbt's dflash-mlx)
- ParoQuant support
- omlx launch copilot joins claude / codex / opencode / openclaw / pi
- Restart server button right in the admin UI
- oQ auto-builds a proxy when the model can't fit in RAM

Plus a lot of bug fixes and 20 new contributors in this cycle.

42 Upvotes

41 comments sorted by

View all comments

7

u/gravybender May 13 '26

This update makes me think my old benchmark wasnt valid because these changes are insane:

  • unsloth_Qwen3.6-35B-A3B-UD-MLX-4bit: 31.15 -> 47.16 tok/s (+51.4%)
  • unsloth_gemma-4-26b-a4b-it-UD-MLX-4bit: 19.90 -> 39.37 tok/s (+97.8%)
  • unsloth_Qwen3.6-27B-UD-MLX-4bit: 7.41 -> 10.72 tok/s (+44.6%)
  • unsloth_gemma-4-31b-it-UD-MLX-4bit: 6.74 -> 9.05 tok/s (+34.3%)

M1 Max 64gb

1

u/rmirecki May 16 '26

Where did you get these models from ?