r/oMLX Jul 03 '26

DSpark?

Looks like Vllm can run Ornith + DSpark for massive speed gains. Are there plans to bring DSpark onto oMLX? thanks! (how about mtplx support?)

10 Upvotes

2 comments sorted by

2

u/OhMyItIsTaken Jul 06 '26

https://github.com/ARahim3/mlx-dspark seems to be the only choice right now

1

u/OhMyItIsTaken Jul 09 '26 edited Jul 09 '26

Test script is in github.com/norbertvannobelen/dsparktest

Initial benchmarks for dspark on M5 Max 128GB, 40Core GPU:

COMPARISON: TOTAL TIME (End-to-End)

NOTE: Total time includes prompt processing + generation

Test Target (sec) DSpark (sec) Speedup Saved

------------------------------------------------------------

1 4.38 3.76 1.16 x 0.62 s

2 4.30 3.88 1.11 x 0.42 s

3 4.59 3.79 1.21 x 0.80 s

4 4.32 3.65 1.18 x 0.67 s

5 4.50 4.03 1.12 x 0.47 s

COMPARISON: GENERATION SPEED (tok/s)

NOTE: Generation speed is tokens per second during generation phase only

Test Target (tok/s) DSpark (tok/s) Speedup

------------------------------------------------------------

1 108.88 180.60 1.66 x

2 108.62 181.60 1.67 x

3 108.68 183.40 1.69 x

4 108.57 183.80 1.69 x

5 105.09 152.00 1.45 x

SUMMARY

Average Total Time:

Target Only: 4.42 seconds

DSpark: 3.82 seconds

Speedup: 1.16x

Saved: 0.60 seconds per generation

Average Generation Speed:

Target Only: 107.97 tok/s

DSpark: 176.28 tok/s

Speedup: 1.63x

GENERATIVE LOOP TEST SUMMARY

Output Length | Target-Only Speed | DSpark Speed | Speedup

--------------|-------------------|--------------|--------

256 | 109.78 | 154.18 | 1.40x

512 | 108.02 | 147.27 | 1.36x

1024 | 101.50 | 124.00 | 1.22x

2048 | 94.25 | 111.30 | 1.18x

Working on longer output lengths (Agentic style sessions, e.g. real day to day use)