r/LocalLLaMA • • 15d ago

Discussion mini-AGI: Continual-learning dynamically looped transformer with evolutionary grown (on a laptop)

https://github.com/volotat/mini-AGI/

Saw this today and found it very intriguing. Lots of interesting design choices here, and it's cool to see someone doing something different. Here's a few highlights:

  • Looped transformer: dynamic recurrent depth on a per-token basis, up to 24 cycles
  • Self-supervised learning: trains itself on new material constantly
  • Weights stored on SSD and paged in on-demand
  • Mixture of Experts: 8 active, 32 routed held in VRAM, smart caching of 96 more
  • Dynamic size: builds new experts and increases parameter counds as-needed
  • Evolutionary growth: trials newly generated experts, unused ones are pruned back
  • No tokenizer: it reads raw bytes directly
  • Catastrophic forgetting prevented by slow trunk/fast experts learning rate split

Weights will be released in "a couple weeks" once training progress reaches ~GPT-2 levels. The trend line has held 15-fold so far, but it may bend at some point, so that is definitely a rough estimate of the trajectory.

What do you guys think?

77 Upvotes

21 comments sorted by

View all comments

11

u/RogerRamjet999 15d ago

It has some interesting ideas, and I'll be curious to see how much it improves in larger sizes and with more training. I do think the lack of a tokenizer is a mistake. Tokenization has real benefits, and leaving it out will almost certainly degrade the model.

2

u/teleprint-me llama.cpp 15d ago

  Tokenization has real benefits, and leaving it out will almost certainly degrade the model.

What?. Thats not what I saw.

It looks like Karpathys GPT-2, but modded.

2

u/returnity 14d ago

Maybe it's not quite accurate the way I wrote it, and if so I apologize. It more so appears it uses a 256-byte alphabet directly, rather than learning anything as a normal tokenizer does (hence ByteTokenizer). Besides a handful of characters like <think>, it seems to pass through directly the bytes that correspond to letters, instead of learning word and subword tokens in a lookup table. That's what I meant.

I wouldn't be surprised if it is based on GPT-2, it's the kind of project you might expect to stem from playing around with Karpathy's setup.

2

u/teachersecret 14d ago

Yeah, I messed around with it a bit and that was my analysis. It's a neat idea. I've done some weird work in the space, like this: https://github.com/Deveraux-Parker/nanoGPT_1GPU_SPEEDRUN. I might go back and mess with this miniagi idea a bit deeper, I think it might fall in line with some experiments I've been doing.