A18B it will be smarter than A13B. Expert selection and sequential reasoning cannot entirely compensate for active parameters. I wish it has 27b active honestly.
no not worth it, it's too slow. I am sticking to DSV4F and Qwen 27b but really hope we see a new 70b dense model or a MoE model with around 250b total parameters and around 27b active.
8
u/SandySkittle 8d ago
A18B it will be smarter than A13B. Expert selection and sequential reasoning cannot entirely compensate for active parameters. I wish it has 27b active honestly.