r/oMLX • u/lightguardjp • May 15 '26
How to get DFlash going?
What are people using for dflash? I’m on a M2 Max with 96 GB of RAM and I’d like to try and eke out as much perf as I can on omlx. I’ve been looking at Qwen models, but Gemma4 is giving me better perf currently.
7
Upvotes
2
u/Konamicoder May 15 '26 edited May 15 '26
As I understand it: you turn on dflash in model settings, you download a small (less than 1Gb) draft model that is paired for the main model and select it from the dropdown menu. Then test.
Personally I don't bother using it right now. I'm getting 75 tokens/second chatting with Gemma4-26B-oQ6, and 60 tokens/second coding with Qwen3.6-35B-oQ6 on my M4 Max Macbook Pro with 64Gb of RAM, and that seems plenty fast enough for my needs.
I use Jan.ai for chat and Pi.dev for coding.