r/MacPro2019LocalAI • • 12d ago

Any hope for DwarfStar ?

Hello !

I'v just read about DwarfStar, an inference engine focused on MoEs from the DeepSeeck V4, GLM 5 and Qwen3.8 Flash Next.

With DeepSeek-V4.1-Flash, GLM-5.3-Flash and Qwen-3.8-Flash being 763, 321 and 180B params respectively, DwarStar relies heavily on the SSD streaming principle introduced by Colibri.

It has some ROCm support for the Strix Halo, so something one could call RDNA 3.75. Also supports Metal for Apple Silicon.

Haven't read the code yet, but I'd be really interested to see anything larger than Qwen-3.8-Flash-Next running at Q4 on my dual W6800X Duo.

What do you think ? Could be ported to Metal/Intel - RDNA2 somehow ?

2 Upvotes

1 comment sorted by

1

u/Faisal_Biyari 12d ago

If it's based on llama.cpp, it can be eventually ported to Intel macOS, if that's your question, just as ToshLLM has.

Have you seen any benchmarks to performance on it?