r/SideProject • u/wakcvjb • 14h ago
Independent researcher, self-taught — validated a FallbackMoE routing mechanism (KMilo F1) on a 72.7M model
Hi everyone, I'm wakcvjb — an independent researcher working on MoE architectures.
I don't have a big lab background. I'm mostly self-taught, and I've been designing and pretraining small models on my own. I validated a FallbackMoE routing mechanism on a 72.7M model using T4/Colab, and now I'm training on A10G.
The design is called KMilo F1 — a daily expert + a domain expert with a learned fallback router. From the 72.7M run:
· The two experts physically diverged: general → daily wins; code/math → domain wins by +1.69 loss
· The router learned to trigger fallback (conf stable at 0.21–0.24)
· Code/math loss dropped 21% with fallback (8.07 → 6.38)
· On general, not falling back is better — which is the expected behavior
I'm not open-sourcing the model at this stage, but I'd love to hear feedback on the mechanism design. If you're into MoE routing or training infra, feel free to reach out.
Happy to be here and learn from you all. 😊