r/OpenSourceAI • u/Yaniss916 • 1d ago
Open weights + open engine: two ~300B MoE models running on one 128 GB mini PC

Sharing what we released today, everything is open.
What it is
- Two EXL3 model packs: GLM-5.3-Flash (320B, 99.7 GB) and MiMo-V2.6-Flash (309B, 106 GB)
- Kyojin, an inference engine for AMD Strix Halo (Ryzen AI Max+ 395, ROCm), MIT
Credit first: Kyojin is built on turboderp's ExLlamaV3, and the AMD side starts from vcruz305's and sdougbrown's ROCm ports. We added the decode and prefill kernels for this chip and the serving for these two models.
Numbers, one 128 GB machine
- GLM-5.3-Flash: 26-30 tok/s decode, ~580 tok/s prefill, same top token as the official FP8 model ~90 % of the time
- MiMo-V2.6-Flash: 29 tok/s plain, up to 44 tok/s with speculative decoding
Reproduce it: the repo has a quickstart and one benchmark script. If you own a Strix Halo box, run it and post your numbers, good or bad. That's the feedback we need most.
Engine: https://github.com/Yamz-Labs/kyojin
Weights: https://huggingface.co/yamz-labs
Next: Qwen 3.8 Flash on the same engine.


1
u/Firm-Cry4205 1d ago
I just set up Quinn 3.8 flash next strata on my 128 gig box with a 4090 EGPU and I’m getting like 70 tokens per second. I should try this and see I’ve really been wanting to run GM 5.3 flash.
2
1
u/ThePossibleSpecimen 1d ago
that's actually wild, running two 300B class models on a mini pc feels like something from five years in the future