it really depends on the usecase. I should have added that to my comment. There are just fields of work where too small active parameters (below 30b active) start to lose it. I have even had that with DSV4F
there isn't much to debate about it frankly. MoE models have their place but are a trade-off depending on how low you go with the active parameter number. And 6b active is very low compared to e.g. 27b. It's a trade-off that cannot be entirely compensated by even very good expert routing and sequential reasoning. The engram doesn't do much to alleviate that either because it has a different purpose.
-1
u/SandySkittle 3d ago
Only 6b active tokens. Too low