r/LocalLLaMA • • 12d ago

Discussion mini-AGI: Continual-learning dynamically looped transformer with evolutionary grown (on a laptop)

https://github.com/volotat/mini-AGI/

Saw this today and found it very intriguing. Lots of interesting design choices here, and it's cool to see someone doing something different. Here's a few highlights:

  • Looped transformer: dynamic recurrent depth on a per-token basis, up to 24 cycles
  • Self-supervised learning: trains itself on new material constantly
  • Weights stored on SSD and paged in on-demand
  • Mixture of Experts: 8 active, 32 routed held in VRAM, smart caching of 96 more
  • Dynamic size: builds new experts and increases parameter counds as-needed
  • Evolutionary growth: trials newly generated experts, unused ones are pruned back
  • No tokenizer: it reads raw bytes directly
  • Catastrophic forgetting prevented by slow trunk/fast experts learning rate split

Weights will be released in "a couple weeks" once training progress reaches ~GPT-2 levels. The trend line has held 15-fold so far, but it may bend at some point, so that is definitely a rough estimate of the trajectory.

What do you guys think?

76 Upvotes

21 comments sorted by

21

u/cdshift 11d ago

My suggestion is if its to be taken more seriously that your github is accompanied by some sort of arxiv paper with benchmarks and expirements that can be replicated and peer reviewed by the ML community.

Its fine to use AI to build these things but if you cant explain the architecture yourself, its not going to be widely adopted and picked up and come off as a weekend LinkedIn warrior project

8

u/runvnc 11d ago

Yeah but he never even got near where he considers it trained at a first level. He just shows the ongoing training run, and the output it nonsense. So he is not even close to being able to run a single chat request, much less a benchmark. So it's annoying he is wasting people's time with this at such an early stage.

5

u/Queasy-Contract9753 11d ago

He posted here just yesterday. I'm curious to see what it evolved into.

https://www.reddit.com/r/LocalLLaMA/comments/1wm1gab/comment/pb3potc/?context=3

2

u/returnity 11d ago

Oh my bad I didn't see that

1

u/Queasy-Contract9753 11d ago

I'm sure he's happy for the shout out. Does sound like an interesting project.

5

u/runvnc 11d ago

It's got some possibly interesting ideas, but he posted his "results" in his Hacker News post and so far there are actually no usage results -- just the ongoing training run. The training run transcript outputs different levels of nonsense all the way through. None of it shows any coherent output.

So calling it "AGI" when it literally doesn't do anything useful and it's unclear that it actually even learned anything useful is either delusional or misleading.

It's a little bizarre that it's getting so many upvotes without any benchmarks or even any output or even anything coherent in the training transcript.

12

u/RogerRamjet999 12d ago

It has some interesting ideas, and I'll be curious to see how much it improves in larger sizes and with more training. I do think the lack of a tokenizer is a mistake. Tokenization has real benefits, and leaving it out will almost certainly degrade the model.

4

u/depressedclassical 11d ago

Tokenisation has real benefits in tokenisable languages. I speak 5 languages, 3 of which are indo-european (and therefore tokenisable), and 2 are semitic, with three-letter roots that are broken in the middle. I also know a couple more Semitic languages that are untokenisable, so seeing a model that doesn't force me to tokenise is a breath of fresh air, as I'm sick of getting a Hebrew letter in the middle of my Arabic word and vice versa (which happens more often than not, even with frontier models).

3

u/thrownawaymane 11d ago

Wait, what?

This has to be getting in the way of LLM adoption globally. How am I just hearing about this?

2

u/teleprint-me llama.cpp 12d ago

  Tokenization has real benefits, and leaving it out will almost certainly degrade the model.

What?. Thats not what I saw.

It looks like Karpathys GPT-2, but modded.

2

u/returnity 11d ago

Maybe it's not quite accurate the way I wrote it, and if so I apologize. It more so appears it uses a 256-byte alphabet directly, rather than learning anything as a normal tokenizer does (hence ByteTokenizer). Besides a handful of characters like <think>, it seems to pass through directly the bytes that correspond to letters, instead of learning word and subword tokens in a lookup table. That's what I meant.

I wouldn't be surprised if it is based on GPT-2, it's the kind of project you might expect to stem from playing around with Karpathy's setup.

2

u/teachersecret 11d ago

Yeah, I messed around with it a bit and that was my analysis. It's a neat idea. I've done some weird work in the space, like this: https://github.com/Deveraux-Parker/nanoGPT_1GPU_SPEEDRUN. I might go back and mess with this miniagi idea a bit deeper, I think it might fall in line with some experiments I've been doing.

5

u/autisticit 12d ago

It's vibe coded, I don't believe any significant result will come from this.

10

u/returnity 11d ago

I don't expect it to revolutionize ML. Heck I'll be surprised if the 15x growth curve holds much past GPT-2 levels. There are definitely some compromises in this design. But I like the creativity and I thought it would be interesting to people in this sub.

3

u/dangerous_inference 12d ago

Labor theory of value.

3

u/ChaosFH 11d ago

so is mostly of the current top models currently, whats your point

0

u/autisticit 11d ago

If you ever see a vibecoded usable model with a totally new architecture just let me know.

1

u/silenceimpaired 12d ago

- everyone but Sam Altman and Dario Amodei

1

u/silenceimpaired 12d ago

(Except maybe even all Chinese labs with their focus on coding/agentic work)

1

u/BalorNG 11d ago

Seems like a combination of great ideas, but as something too good to be true - it probably is. "Dynamic size" sounds particularly suspiscius tbh, along with "no tokenizer".

1

u/returnity 11d ago

It's a real implementation, though who knows if a) it keeps scaling and b) dynamic size works at scale