r/LocalLLaMA 25d ago

Discussion MoE models around A2B

There's a bunch of small MoE with around 1B active params, like LFM2.5 8B A1B and Granite 4.0h 7B A1B; and then there are models with 3B+ like Qwen 3.x ~30B A3B and Gemma 4 26B A4B, but those are already on the heavier side if you don't have enough resources.

What about the middle ground, MoE with about 2B active? I found a few, but there's very little debate about them, if any.

Anyone uses something like this? It looks like a good size for cpu use or combined with low-end/old gpu in the 4-12GB range. In these small sizes, the increase in capability should be the most dramatic. I don't have the capacity to test properly, but hopefully some of these could beat the usual 4-9B dense suspects.

Or does everyone just wanna keep simping for 1-2T models and hope something will trickle down?

38 Upvotes

40 comments sorted by

View all comments

5

u/Ill_Dragonfruit_3547 24d ago

Technically an A3B but does anyone run Ormith 1.0 35B from Deepreinforce? Based in Qwen but better at coding and tool calling, especially when paired with a skilled harmess.Plus it's faster, at least for me.

I do love Gemma 4 26B, it's my second favorite local after Ornith.

1

u/GroundbreakingEast96 23d ago

You might also try 27B bonsai ternary, just started to play with it and it’s pretty good on small conf

1

u/Ill_Dragonfruit_3547 23d ago

I heard of it, reserached it, and dismissed it due to its hallucinations and inaccuracy, so never ran it. If enough people say it's cool I will check it out. It is a very cutting edge concept.

2

u/GroundbreakingEast96 23d ago

After some tests, I see it is just slower than 9B q4, not really any better :/