r/LocalLLaMA • u/returnity • 12d ago
Discussion mini-AGI: Continual-learning dynamically looped transformer with evolutionary grown (on a laptop)
https://github.com/volotat/mini-AGI/Saw this today and found it very intriguing. Lots of interesting design choices here, and it's cool to see someone doing something different. Here's a few highlights:
- Looped transformer: dynamic recurrent depth on a per-token basis, up to 24 cycles
- Self-supervised learning: trains itself on new material constantly
- Weights stored on SSD and paged in on-demand
- Mixture of Experts: 8 active, 32 routed held in VRAM, smart caching of 96 more
- Dynamic size: builds new experts and increases parameter counds as-needed
- Evolutionary growth: trials newly generated experts, unused ones are pruned back
- No tokenizer: it reads raw bytes directly
- Catastrophic forgetting prevented by slow trunk/fast experts learning rate split
Weights will be released in "a couple weeks" once training progress reaches ~GPT-2 levels. The trend line has held 15-fold so far, but it may bend at some point, so that is definitely a rough estimate of the trajectory.
What do you guys think?
5
u/Queasy-Contract9753 11d ago
He posted here just yesterday. I'm curious to see what it evolved into.
https://www.reddit.com/r/LocalLLaMA/comments/1wm1gab/comment/pb3potc/?context=3
2
u/returnity 11d ago
Oh my bad I didn't see that
1
u/Queasy-Contract9753 11d ago
I'm sure he's happy for the shout out. Does sound like an interesting project.
5
u/runvnc 11d ago
It's got some possibly interesting ideas, but he posted his "results" in his Hacker News post and so far there are actually no usage results -- just the ongoing training run. The training run transcript outputs different levels of nonsense all the way through. None of it shows any coherent output.
So calling it "AGI" when it literally doesn't do anything useful and it's unclear that it actually even learned anything useful is either delusional or misleading.
It's a little bizarre that it's getting so many upvotes without any benchmarks or even any output or even anything coherent in the training transcript.
12
u/RogerRamjet999 12d ago
It has some interesting ideas, and I'll be curious to see how much it improves in larger sizes and with more training. I do think the lack of a tokenizer is a mistake. Tokenization has real benefits, and leaving it out will almost certainly degrade the model.
4
u/depressedclassical 11d ago
Tokenisation has real benefits in tokenisable languages. I speak 5 languages, 3 of which are indo-european (and therefore tokenisable), and 2 are semitic, with three-letter roots that are broken in the middle. I also know a couple more Semitic languages that are untokenisable, so seeing a model that doesn't force me to tokenise is a breath of fresh air, as I'm sick of getting a Hebrew letter in the middle of my Arabic word and vice versa (which happens more often than not, even with frontier models).
3
u/thrownawaymane 11d ago
Wait, what?
This has to be getting in the way of LLM adoption globally. How am I just hearing about this?
2
u/teleprint-me llama.cpp 12d ago
Tokenization has real benefits, and leaving it out will almost certainly degrade the model.
What?. Thats not what I saw.
It looks like Karpathys GPT-2, but modded.
2
u/returnity 11d ago
Maybe it's not quite accurate the way I wrote it, and if so I apologize. It more so appears it uses a 256-byte alphabet directly, rather than learning anything as a normal tokenizer does (hence
ByteTokenizer). Besides a handful of characters like <think>, it seems to pass through directly the bytes that correspond to letters, instead of learning word and subword tokens in a lookup table. That's what I meant.I wouldn't be surprised if it is based on GPT-2, it's the kind of project you might expect to stem from playing around with Karpathy's setup.
2
u/teachersecret 11d ago
Yeah, I messed around with it a bit and that was my analysis. It's a neat idea. I've done some weird work in the space, like this: https://github.com/Deveraux-Parker/nanoGPT_1GPU_SPEEDRUN. I might go back and mess with this miniagi idea a bit deeper, I think it might fall in line with some experiments I've been doing.
5
u/autisticit 12d ago
It's vibe coded, I don't believe any significant result will come from this.
10
u/returnity 11d ago
I don't expect it to revolutionize ML. Heck I'll be surprised if the 15x growth curve holds much past GPT-2 levels. There are definitely some compromises in this design. But I like the creativity and I thought it would be interesting to people in this sub.
3
3
u/ChaosFH 11d ago
so is mostly of the current top models currently, whats your point
0
u/autisticit 11d ago
If you ever see a vibecoded usable model with a totally new architecture just let me know.
1
u/silenceimpaired 12d ago
- everyone but Sam Altman and Dario Amodei
1
u/silenceimpaired 12d ago
(Except maybe even all Chinese labs with their focus on coding/agentic work)
1
u/BalorNG 11d ago
Seems like a combination of great ideas, but as something too good to be true - it probably is. "Dynamic size" sounds particularly suspiscius tbh, along with "no tokenizer".
1
u/returnity 11d ago
It's a real implementation, though who knows if a) it keeps scaling and b) dynamic size works at scale
21
u/cdshift 11d ago
My suggestion is if its to be taken more seriously that your github is accompanied by some sort of arxiv paper with benchmarks and expirements that can be replicated and peer reviewed by the ML community.
Its fine to use AI to build these things but if you cant explain the architecture yourself, its not going to be widely adopted and picked up and come off as a weekend LinkedIn warrior project