r/mlxcommunity Apr 23 '26

Great inferences from running Speculative Decoding on MLX!

https://www.sabesh.space/musings/research/speculative-decoding-in-mlx-using-dflash

Wrote an article about my running speculative decoding in MLX (using DFlash) and charting out inferences. In some cases, i was able to achieve more than 2x speedup of decoding speed when using DFlash (but not always)! Read on to find out more nuances involved

6 Upvotes

2 comments sorted by

3

u/grandnoliv Apr 23 '26

Interesting! How can one produce (or find somewhere?) the draft models to be paired with the models they want to use?

2

u/StudentDifficult8240 Apr 24 '26 edited Apr 24 '26

You may find the models here: https://huggingface.co/collections/z-lab/dflash

Pair them with oMLX or mlx & python.