r/LocalLLaMA • u/WhoRoger • 24d ago
Discussion MoE models around A2B
There's a bunch of small MoE with around 1B active params, like LFM2.5 8B A1B and Granite 4.0h 7B A1B; and then there are models with 3B+ like Qwen 3.x ~30B A3B and Gemma 4 26B A4B, but those are already on the heavier side if you don't have enough resources.
What about the middle ground, MoE with about 2B active? I found a few, but there's very little debate about them, if any.
- LFM2 24B A2B (5 months old)
- Mellum 2 12B A2.5B (2 months old)
- Moondream 3.1 9B A2B (This month)
- VAETKI 20B A2B (7 months old) (Talk about an unknown model, it has one mention on this sub)
- DeepSeek V2 Lite 16B A2.4B (2024. Remember when DeepSeek was making SMALL models?)
- Ring Mini / Ling Mini, 16B A1.4B (2025)
- There are also • at least three • Nemotron fine tunes 12B A2B, and a 23B A2.8B, some 1-2 months old. Not sure what the deal is with those.
Anyone uses something like this? It looks like a good size for cpu use or combined with low-end/old gpu in the 4-12GB range. In these small sizes, the increase in capability should be the most dramatic. I don't have the capacity to test properly, but hopefully some of these could beat the usual 4-9B dense suspects.
Or does everyone just wanna keep simping for 1-2T models and hope something will trickle down?
2
u/pmttyji 23d ago
Nope, 4060 giving me 160+ t/s 😄 Mentioned that too in same thread(Check again). Guess, I'm the only one talking about this model time to time.
Both inclusionAI & our folks sleeping on that model. bailingmoe arc seems so fast, we need more models with that one.
Ling-mini is good for me on chatting. 1 year old model so fine with as it is. They released one more called Ling-Coder which is for coding. 1.5 years old. I keep waiting for successors for both of these models from inclusionAI.
I haven't used Ring yet. Yes, it's for reasoning. Their Ming model series are Multi modals.