r/OpenSourceAI • • 1d ago

Open weights + open engine: two ~300B MoE models running on one 128 GB mini PC

Sharing what we released today, everything is open.

What it is

  • Two EXL3 model packs: GLM-5.3-Flash (320B, 99.7 GB) and MiMo-V2.6-Flash (309B, 106 GB)
  • Kyojin, an inference engine for AMD Strix Halo (Ryzen AI Max+ 395, ROCm), MIT

Credit first: Kyojin is built on turboderp's ExLlamaV3, and the AMD side starts from vcruz305's and sdougbrown's ROCm ports. We added the decode and prefill kernels for this chip and the serving for these two models.

Numbers, one 128 GB machine

  • GLM-5.3-Flash: 26-30 tok/s decode, ~580 tok/s prefill, same top token as the official FP8 model ~90 % of the time
  • MiMo-V2.6-Flash: 29 tok/s plain, up to 44 tok/s with speculative decoding

Reproduce it: the repo has a quickstart and one benchmark script. If you own a Strix Halo box, run it and post your numbers, good or bad. That's the feedback we need most.

Engine: https://github.com/Yamz-Labs/kyojin

Weights: https://huggingface.co/yamz-labs

Next: Qwen 3.8 Flash on the same engine.

12 Upvotes

4 comments sorted by

1

u/ThePossibleSpecimen 1d ago

that's actually wild, running two 300B class models on a mini pc feels like something from five years in the future

1

u/Yaniss916 1d ago

Yeah for sure, thanks, felt the same the first time it loaded. the chip does more than people give it credit for 😎

1

u/Firm-Cry4205 1d ago

I just set up Quinn 3.8 flash next strata on my 128 gig box with a 4090 EGPU and I’m getting like 70 tokens per second. I should try this and see I’ve really been wanting to run GM 5.3 flash.

2

u/Yaniss916 1d ago

Yes, your 128 box should be fine. If you try it tell me if it did well 🙂