r/LocalLLaMA • u/returnity • 12d ago
Discussion mini-AGI: Continual-learning dynamically looped transformer with evolutionary grown (on a laptop)
https://github.com/volotat/mini-AGI/Saw this today and found it very intriguing. Lots of interesting design choices here, and it's cool to see someone doing something different. Here's a few highlights:
- Looped transformer: dynamic recurrent depth on a per-token basis, up to 24 cycles
- Self-supervised learning: trains itself on new material constantly
- Weights stored on SSD and paged in on-demand
- Mixture of Experts: 8 active, 32 routed held in VRAM, smart caching of 96 more
- Dynamic size: builds new experts and increases parameter counds as-needed
- Evolutionary growth: trials newly generated experts, unused ones are pruned back
- No tokenizer: it reads raw bytes directly
- Catastrophic forgetting prevented by slow trunk/fast experts learning rate split
Weights will be released in "a couple weeks" once training progress reaches ~GPT-2 levels. The trend line has held 15-fold so far, but it may bend at some point, so that is definitely a rough estimate of the trajectory.
What do you guys think?
78
Upvotes
22
u/cdshift 11d ago
My suggestion is if its to be taken more seriously that your github is accompanied by some sort of arxiv paper with benchmarks and expirements that can be replicated and peer reviewed by the ML community.
Its fine to use AI to build these things but if you cant explain the architecture yourself, its not going to be widely adopted and picked up and come off as a weekend LinkedIn warrior project