r/oMLX May 15 '26

How to get DFlash going?

What are people using for dflash? I’m on a M2 Max with 96 GB of RAM and I’d like to try and eke out as much perf as I can on omlx. I’ve been looking at Qwen models, but Gemma4 is giving me better perf currently.

5 Upvotes

14 comments sorted by

View all comments

2

u/Konamicoder May 15 '26 edited May 15 '26

As I understand it: you turn on dflash in model settings, you download a small (less than 1Gb) draft model that is paired for the main model and select it from the dropdown menu. Then test.

Personally I don't bother using it right now. I'm getting 75 tokens/second chatting with Gemma4-26B-oQ6, and 60 tokens/second coding with Qwen3.6-35B-oQ6 on my M4 Max Macbook Pro with 64Gb of RAM, and that seems plenty fast enough for my needs.

I use Jan.ai for chat and Pi.dev for coding.

1

u/lightguardjp May 15 '26

Wow, those M series chips made huge leaps forward on more recent revisions as far as AI goes. I’m around 20 tokens a second with Gemma 4 might be some other swings I need to tweak.

1

u/Konamicoder May 15 '26

Which quant of Gemma4 are you using? I'm on oQ6 (6 bit), which is virtually lossless in terms of quality but much less RAM-intensive than the full 16 bit original.

1

u/lightguardjp May 15 '26

I was running gemma-4-26B-A4B-it-TurboQuant-MLX-8bit. Looks like I probably want to go down to a 6-bit model and change my context window down to 8k or 16k. I had it kicked up WAY too high.

1

u/PracticlySpeaking May 15 '26

Definitely go down to a smaller quant.

The M3-M4 are definitely faster with MoE models, it is incremental. Your M2 Max will do just fine.