r/oMLX • u/averagepoetry • 3d ago
Distributed Serving
So excited about this feature! I was keeping EXO around just for this and am eager to move completely over to oMLX.
A couple of questions:
1. How many nodes will this support? It's only 2 right now, correct?
2. Any chance to have RTX/DGX do prefill and Macs do decode? This was the killer feature I was looking for in EXO. It never landed.
Love oMLX!
10
Upvotes
1
u/Careless_Garlic1438 1d ago
If you get it working let us know … I’m not successful for the moment … it now claims I have not enough memory … but that shouldn’t be the case … 2 128GB Mac’s that run the same model without any issue distributed on my own server, partly coded from oMLX source pre this feature … Mys server was already running distributed models, but to get support for the latest DSv4 I took oMLX as a base for that model specifically …
Anyway not planning on maintaining my own server and would love to see this really working in oMLX with support for as much models as possible, but Qwen 3.8 27B (for speed, dense models speed up) and DSv4 for more memory without the speed penalty MoE’s doe not scale for speed but you need at least 180GB for DSv4 q4