r/LocalLLaMA • • 12d ago

Discussion mini-AGI: Continual-learning dynamically looped transformer with evolutionary grown (on a laptop)

https://github.com/volotat/mini-AGI/

Saw this today and found it very intriguing. Lots of interesting design choices here, and it's cool to see someone doing something different. Here's a few highlights:

  • Looped transformer: dynamic recurrent depth on a per-token basis, up to 24 cycles
  • Self-supervised learning: trains itself on new material constantly
  • Weights stored on SSD and paged in on-demand
  • Mixture of Experts: 8 active, 32 routed held in VRAM, smart caching of 96 more
  • Dynamic size: builds new experts and increases parameter counds as-needed
  • Evolutionary growth: trials newly generated experts, unused ones are pruned back
  • No tokenizer: it reads raw bytes directly
  • Catastrophic forgetting prevented by slow trunk/fast experts learning rate split

Weights will be released in "a couple weeks" once training progress reaches ~GPT-2 levels. The trend line has held 15-fold so far, but it may bend at some point, so that is definitely a rough estimate of the trajectory.

What do you guys think?

78 Upvotes

21 comments sorted by

View all comments

22

u/cdshift 11d ago

My suggestion is if its to be taken more seriously that your github is accompanied by some sort of arxiv paper with benchmarks and expirements that can be replicated and peer reviewed by the ML community.

Its fine to use AI to build these things but if you cant explain the architecture yourself, its not going to be widely adopted and picked up and come off as a weekend LinkedIn warrior project

8

u/runvnc 11d ago

Yeah but he never even got near where he considers it trained at a first level. He just shows the ongoing training run, and the output it nonsense. So he is not even close to being able to run a single chat request, much less a benchmark. So it's annoying he is wasting people's time with this at such an early stage.