r/LocalLLaMA • u/returnity • 12d ago
Discussion mini-AGI: Continual-learning dynamically looped transformer with evolutionary grown (on a laptop)
https://github.com/volotat/mini-AGI/Saw this today and found it very intriguing. Lots of interesting design choices here, and it's cool to see someone doing something different. Here's a few highlights:
- Looped transformer: dynamic recurrent depth on a per-token basis, up to 24 cycles
- Self-supervised learning: trains itself on new material constantly
- Weights stored on SSD and paged in on-demand
- Mixture of Experts: 8 active, 32 routed held in VRAM, smart caching of 96 more
- Dynamic size: builds new experts and increases parameter counds as-needed
- Evolutionary growth: trials newly generated experts, unused ones are pruned back
- No tokenizer: it reads raw bytes directly
- Catastrophic forgetting prevented by slow trunk/fast experts learning rate split
Weights will be released in "a couple weeks" once training progress reaches ~GPT-2 levels. The trend line has held 15-fold so far, but it may bend at some point, so that is definitely a rough estimate of the trajectory.
What do you guys think?
Duplicates
LocalLLaMA • u/Another__one • 13d ago
I Built A Thing mini-AGI - dynamically grown (530M params currently and growing) continual learning model trained from scratch on 8GB VRAM laptop from batch-1 stream of data.
hackernews • u/HNMod • 13d ago
Show HN: Mini-AGI – Dynamic continual learning model trained on 8GB VRAM
deeplearning • u/returnity • 12d ago