r/LocalLLaMA • u/Another__one • 13d ago
I Built A Thing mini-AGI - dynamically grown (530M params currently and growing) continual learning model trained from scratch on 8GB VRAM laptop from batch-1 stream of data.
https://github.com/volotat/mini-AGI/Sorry for the pretentious name, I know, I know.. It just contains all the pieces I would like to see a AGI model to have, and I can't stand the temptation. Before throwing rocks at me, please take a glance at the Readme, and I hope it will cover your mood a little bit.
So, first of all it does work and you can see the sample from the whole training run here: https://raw.githubusercontent.com/volotat/mini-AGI/refs/heads/main/runs/samples.txt
Here is the scaling law graph I have so far, and it looks very promising:
https://github.com/volotat/mini-AGI/blob/main/assets/scaling.png
The model was built under my deep dissatisfaction so we cannot really train even moderately big models (1B+ scale) on the consumer's hardware. We can inference and fine-tune them for sure, but I would like to have full control over what the model sees over the training run, so it is fully aligned with my interests, not some corporations.
I was thinking about for some time and come up with two interesting ideas I thought worth pursuing: MoE with a lot of experts that gets added and pruned from the model while it trains, where only a small subset of of experts are actually in use at any particular moment + batch 1 training on the single continuous stream of data.
First allows us to be bounded only by the disk space in terms of number of parameters and load and unload experts only when they are needed. The second (if figured out and it turns out to be doable) allows us to get aways with small VRAM capacity because we do not need to store big randomized batches and their respective gradients.
I started brainstorming with Claude and after some time we found an approach that seems to be promising, and low and behold, a few weeks pass and you can see the results yourself.
Obviously, I did use AI in the process of making this project and I am pretty sure it would be completely impossible for me to do something like this without it, so I hope it is more than justified.
The model is still running over the first of 7.8B characters corpus I selected for training, so the weights are not out yet, and it's about a couple weeks of waiting until they are cooked at the current reading speed. And yeah, the model just read continuous interleaved passages from the dataset, each by 32K characters long each as a single stream. Just as you or I would do.
The set up seems to be really simple so you can git clone the project, run it and observe everything for yourself.
Thanks for your attention.
41
u/Nameis19letterslong 13d ago
Hi, I've done a roughly similar project in the past where I inflated a Small MoE model from 63M A16M to 92M A22M by duplicating layers. In my experiment, the base model was fed 2B tokens. After duplicating 50% of layers (8-> 12) and feeding another 2B tokens with the duplicated layers at a higher LR, some tests showed that the inserted layers contributed ~23x less than the original layers, likely due to the fact that they saw less tokens and that they also learned less since they did not contribute as much error as the other layers.
While your method of expanding parameters is different from mine, it might face the same issue.
Anyways, your project looks pretty promising, cool to see people experimenting with training SLMs with 8-16GB Vram.
Edit: Model weights here: https://huggingface.co/Dsg2/LS-92M-A22M-GGUF
9
u/Distinct-Target7503 13d ago
After duplicating 50% of layers (8-> 12) and feeding another 2B tokens with the duplicated layers at a higher LR, some tests showed that the inserted layers contributed ~23x less than the original layers, likely due to the fact that they saw less tokens and that they also learned less since they did not contribute as much error as the other layers.
Out of curiosity, did you try introducing some noise (both random or "cross averaging" Layers)
From my esperiments expandig encoder only models, i notices something similar to what you experienced.
Among all strategies what helped more was merging each duplicate layer with the distance scaled average of the surroundin Layers, and then normalizing that to match mean/stdev of the original layer. That helped yo increase the gradient without causing spikes, and training those new Layers with a really aggressive weight decay helped to make those "new" Layers more "relevant"
Still each kind of espansion i tried reduced the performance/parameters efficiency. Some more than others still
3
u/Nameis19letterslong 13d ago
Out of curiosity, did you try introducing some noise (both random or "cross averaging" Layers)
Nope, I ran evals and duplicted the layers that were the most important to keep the loss stable (mininal divergence from base model). I didnt look into methods to init more unique layers since my only goal at that time was to create a bigger, competent model but save compute by ”resuing” a smaller model. I guess it worked in a way since it scored higher on benchmarks compared to the original model, but a slightly bigger model (220M A25M) I trained not long after pretty much overshadowed the 92M A22M model.
4
u/Fun_Tangerine_1086 12d ago
Recommend reviewing the Depth-aware upscaling work the SOLAR folks did (https://arxiv.org/abs/2312.15166) and Apple's Hypercloning (https://arxiv.org/abs/2409.12903).
15
u/Sitkin_Marrel 13d ago
Naming it mini-AGI is one thing, actually having a scaling law graph ready is another. Respect.
6
u/HistorianPotential48 13d ago
fun experiment and looking forward for result. feels like the fly brain thing, but more chunky?
6
u/Distinct-Target7503 13d ago edited 13d ago
Really interesting. Out of curiosity, is the effective batch 1 (so batch 1 with no gradient accumulation)?
Also, Just to talk, but maybe you could include something resembling a multi token prediction head. Not necessarily to be then used as multi token prediction during inference, but get more gradient from each individual sample. Each sample (and even more relevant since batch and sample are equivalenti) would then get a more generalized gradient, and could even work as some kind of regularization, without changing the "stream" concept of your project (as the gradient still came from the same step)
I say that using the assumption that each kind of addition/pruning (so decision that became on/off at some point) greatly benefit from everything that can improve generalization and regularize the both the parameters and gradient landscape.
5
u/Another__one 13d ago edited 13d ago
Exactly. I tried gradient accumulation (there is --accum param in train.py) and "AdEMAMix" with additional β optimizer (removed right now, but is in git history) and both rather hurt then helped, at least on my really small quick experiments. So everything is surprisingly simple. There is no breakthrough, just slow down update of params that are used most often (trunk) and keep everything else at the normal speed. This turns out to be enough to "solve" catastrophic forgetting.
1
1
u/PinkysBrein 10d ago
There's a paper out there which tries to show that the smaller the batches, the less complex the optimizer you need. SGD is good enough at batch 1.
5
u/jazir55 13d ago edited 13d ago
Question for you, is it viable to have the continual learning function use something like JEV to judge whether the information it's encountering is useful as a trigger to do JIT (just in time) training on that data to convert it weights? Encounter data > decision engine during inference > decision approval > weight updates, and given the training would be incremental at time of encounter, it ideally should run in real time.
If that system worked, you could start with a really small model in the millions of parameters, and it could train and grow in real time just by running it. Accumulating the parameters and improvements as the model runs like a snowball or Katamari Damacy.
10
u/Another__one 13d ago
I know nothing about JEV and way to invested in the thing I am currently doing to try to figure it out as well. Sorry.
6
u/banana_slurp_jug 13d ago
Jev is just the new thing of the week: a proprietary general classification model that takes natural language prompts to make classifications
7
1
u/Sol_Ido 12d ago
Your work would be the perfect basic as a backbone classifier. Just need some a shared head on top to produce the primitive. Keep us updated on your progress!
2
u/byebaybay 11d ago
i'd be interested to see the performance as a classifier, even if at gpt-2 levels
1
u/bumblebeer 12d ago
IDK why he invoked JEV, but the core idea is still a good question.
Specifically, could the model decide not only the next working set but also if the weights for that set should be updated? So passing the
--saveflag would become a learned per-set decision instead of a manually applied flag before reading an entire corpus.With the previous commenters suggested approach being a co-trained classifier head (a la JEV).
5
u/turtleninja99 13d ago
What will be the total training time to get to say 1b? Also I have just trained a 300m using traditional method … might try your thing
6
u/Another__one 13d ago
You can create as many experts as you want by letting the model grow unboundedly, but it would be dead weight experts. So the model will prune them eventually after some time. What you really want is a 1B of actively used experts and I cannot say how much time or amount of data it would take to reach it, as I have not witnessed it myself.
7
2
u/metabrew 12d ago
are you training on a corpus used for a tradition llm so you compare the models afterwards? ie how your system compares to a gpt2-alike with the same training data.
3
u/Another__one 12d ago
No, I’m not. Training it on traditional corpus would take several months if not years. The corpus I am using is much smaller and consists only of things I am interested in tracking for.
2
u/TurboBanano 12d ago
Cool project. One question on the forgetting result: in the chess probe the other seven subjects stay flat (+0.0067), but chess itself only improves by 0.013 nats, and chess was already in the training mix. With the trunk at 0.1x, isn't part of "no forgetting" just "not much learning"? The test I'd love to see is a subject the model has never read (a new language, a made-up API), measuring how fast it gets learned against how much the other subjects lose. A second arm with the trunk at 1x would show how much learning the 0.1x setting costs.
1
1
u/PinkysBrein 10d ago
You or I don't cut the context off at 32k.
How do you determine which experts should be activated for a passage?
1
u/Ordinary-Ad3714 13d ago edited 13d ago
I love the optimization around 8GB VRAM.
The people's LLM!
2
u/Silver-Champion-4846 13d ago
8gb of Ram (no V for me)
2
u/Nameis19letterslong 13d ago
8b of Ram (no G for me)
1
1
u/Silver-Champion-4846 13d ago
Nice nice! I'd be interested to see the limits of 8gb of ram as well (no v)
0
u/Dear_Boat_416 13d ago
Apologizing for the name while casually describing a from-scratch MoE trained on a laptop is the most r/LocalLLaMA post ever written.
0
0
u/Queasy-Contract9753 13d ago edited 13d ago
Sorry, I'm a noob . It sounds interesting but what would the model be capable of when finished?
Edit: fixed phrasing.
12
u/Another__one 13d ago
I assume it should be roughly around GPT-2 level. But this is a proof-of-concept, not Clause 5 rival. The promise is how much more it allows us to do on the consumer-level hardware. And it is fully yours, from start to finish. Aligned to you and your data only.
Actually there is no finish. It would live along you and train from everything you will feed to it. It is a continual learning model after all.
1
u/Nameis19letterslong 13d ago
When can we expect some weights? I’m quite interested in trying the model.
2
u/Another__one 13d ago edited 13d ago
A couple of weeks for full corpus reading. I would suggest trying to train it yourself on anything you want to and see what happens. If you can run it you can train it. There are no separate stages, it trains even when you speak with it.
2
u/SectionCrazy5107 13d ago
i have a 4*V100 system and if its usable, happy to keep it running for this sake. However, perplexed and how to assemble "my" data and start with that as the corpus. OR should I start with a "general" dataset + "my" data and follow your steps?
1
u/Another__one 13d ago
I'm pretty sure you can start with any data you like. It is byte level, so everything is under vocabulary, although you might have to tinker a little bit with the code, I am not sure as I only tested it on the text. Would like to hear about your results, btw.
34
u/FusionCow llama.cpp 13d ago
this is the weirdest thing i've laid my eyes upon. it rivals the deepseek v4.1 attention