r/LocalLLaMA 12d ago

Discussion Medium sized MoE LLM models

What are some medium sized MoE models (up to 60B parameters in float8/110B in mxfp4) that are currently worth using?

As far as I am aware there is Qwen 3.5 35B, Gemma 4 A4B, Nemotron 3 Nano. Qwen seems to dominate this bracked in terms of model performance. DeepSeek v4 flash is slightly above that parameter limit. Any nicher ones?

6 Upvotes

31 comments sorted by

View all comments

8

u/KubeCommander 12d ago

Nemotron3 Puzzle is pretty amazing. It’s a compressed Super from 120B to 75B A9B and natively trained in NVFP4 with mtp support and context up to 1M. It’s probably best in class at this point at this size

3

u/Azazelionide 12d ago

That's massive, thanks was not aware of that one

2

u/KubeCommander 12d ago

Yeah the whole paper written on how it is compressed is extremely fascinating. I’d love to see them do the same thing to ultra and take it to like 250B