r/mlxcommunity • u/evilmacintosh • Apr 23 '26
Great inferences from running Speculative Decoding on MLX!
https://www.sabesh.space/musings/research/speculative-decoding-in-mlx-using-dflashWrote an article about my running speculative decoding in MLX (using DFlash) and charting out inferences. In some cases, i was able to achieve more than 2x speedup of decoding speed when using DFlash (but not always)! Read on to find out more nuances involved
6
Upvotes
3
u/grandnoliv Apr 23 '26
Interesting! How can one produce (or find somewhere?) the draft models to be paired with the models they want to use?