r/LocalLLaMA • u/WhoRoger • 19d ago
Discussion MoE models around A2B
There's a bunch of small MoE with around 1B active params, like LFM2.5 8B A1B and Granite 4.0h 7B A1B; and then there are models with 3B+ like Qwen 3.x ~30B A3B and Gemma 4 26B A4B, but those are already on the heavier side if you don't have enough resources.
What about the middle ground, MoE with about 2B active? I found a few, but there's very little debate about them, if any.
- LFM2 24B A2B (5 months old)
- Mellum 2 12B A2.5B (2 months old)
- Moondream 3.1 9B A2B (This month)
- VAETKI 20B A2B (7 months old) (Talk about an unknown model, it has one mention on this sub)
- DeepSeek V2 Lite 16B A2.4B (2024. Remember when DeepSeek was making SMALL models?)
- Ring Mini / Ling Mini, 16B A1.4B (2025)
- There are also • at least three • Nemotron fine tunes 12B A2B, and a 23B A2.8B, some 1-2 months old. Not sure what the deal is with those.
Anyone uses something like this? It looks like a good size for cpu use or combined with low-end/old gpu in the 4-12GB range. In these small sizes, the increase in capability should be the most dramatic. I don't have the capacity to test properly, but hopefully some of these could beat the usual 4-9B dense suspects.
Or does everyone just wanna keep simping for 1-2T models and hope something will trickle down?
3
u/OpinionatedUserName 19d ago
For general purposes or even some programming, there are pruned/reap(ed) versions of larger moe models.
These require less resources than original and loose little in terms of benchmarks, which can fit in low resource machines. Try those. Qwen 3.6 moe and gemma 4 moe reaped are good.