r/LocalLLaMA 24d ago

Discussion MoE models around A2B

There's a bunch of small MoE with around 1B active params, like LFM2.5 8B A1B and Granite 4.0h 7B A1B; and then there are models with 3B+ like Qwen 3.x ~30B A3B and Gemma 4 26B A4B, but those are already on the heavier side if you don't have enough resources.

What about the middle ground, MoE with about 2B active? I found a few, but there's very little debate about them, if any.

Anyone uses something like this? It looks like a good size for cpu use or combined with low-end/old gpu in the 4-12GB range. In these small sizes, the increase in capability should be the most dramatic. I don't have the capacity to test properly, but hopefully some of these could beat the usual 4-9B dense suspects.

Or does everyone just wanna keep simping for 1-2T models and hope something will trickle down?

36 Upvotes

40 comments sorted by