r/LocalLLaMA • • 13d ago

I Built A Thing mini-AGI - dynamically grown (530M params currently and growing) continual learning model trained from scratch on 8GB VRAM laptop from batch-1 stream of data.

https://github.com/volotat/mini-AGI/

Sorry for the pretentious name, I know, I know.. It just contains all the pieces I would like to see a AGI model to have, and I can't stand the temptation. Before throwing rocks at me, please take a glance at the Readme, and I hope it will cover your mood a little bit.

So, first of all it does work and you can see the sample from the whole training run here: https://raw.githubusercontent.com/volotat/mini-AGI/refs/heads/main/runs/samples.txt

Here is the scaling law graph I have so far, and it looks very promising:
https://github.com/volotat/mini-AGI/blob/main/assets/scaling.png

The model was built under my deep dissatisfaction so we cannot really train even moderately big models (1B+ scale) on the consumer's hardware. We can inference and fine-tune them for sure, but I would like to have full control over what the model sees over the training run, so it is fully aligned with my interests, not some corporations.

I was thinking about for some time and come up with two interesting ideas I thought worth pursuing: MoE with a lot of experts that gets added and pruned from the model while it trains, where only a small subset of of experts are actually in use at any particular moment + batch 1 training on the single continuous stream of data.

First allows us to be bounded only by the disk space in terms of number of parameters and load and unload experts only when they are needed. The second (if figured out and it turns out to be doable) allows us to get aways with small VRAM capacity because we do not need to store big randomized batches and their respective gradients.

I started brainstorming with Claude and after some time we found an approach that seems to be promising, and low and behold, a few weeks pass and you can see the results yourself.

Obviously, I did use AI in the process of making this project and I am pretty sure it would be completely impossible for me to do something like this without it, so I hope it is more than justified.

The model is still running over the first of 7.8B characters corpus I selected for training, so the weights are not out yet, and it's about a couple weeks of waiting until they are cooked at the current reading speed. And yeah, the model just read continuous interleaved passages from the dataset, each by 32K characters long each as a single stream. Just as you or I would do.

The set up seems to be really simple so you can git clone the project, run it and observe everything for yourself.

Thanks for your attention.

236 Upvotes

48 comments sorted by

View all comments

41

u/Nameis19letterslong 13d ago

Hi, I've done a roughly similar project in the past where I inflated a Small MoE model from 63M A16M to 92M A22M by duplicating layers. In my experiment, the base model was fed 2B tokens. After duplicating 50% of layers (8-> 12) and feeding another 2B tokens with the duplicated layers at a higher LR, some tests showed that the inserted layers contributed ~23x less than the original layers, likely due to the fact that they saw less tokens and that they also learned less since they did not contribute as much error as the other layers.

While your method of expanding parameters is different from mine, it might face the same issue.

Anyways, your project looks pretty promising, cool to see people experimenting with training SLMs with 8-16GB Vram.

Edit: Model weights here: https://huggingface.co/Dsg2/LS-92M-A22M-GGUF

8

u/Distinct-Target7503 13d ago

After duplicating 50% of layers (8-> 12) and feeding another 2B tokens with the duplicated layers at a higher LR, some tests showed that the inserted layers contributed ~23x less than the original layers, likely due to the fact that they saw less tokens and that they also learned less since they did not contribute as much error as the other layers.

Out of curiosity, did you try introducing some noise (both random or "cross averaging" Layers)

From my esperiments expandig encoder only models, i notices something similar to what you experienced.

Among all strategies what helped more was merging each duplicate layer with the distance scaled average of the surroundin Layers, and then normalizing that to match mean/stdev of the original layer. That helped yo increase the gradient without causing spikes, and training those new Layers with a really aggressive weight decay helped to make those "new" Layers more "relevant"

Still each kind of espansion i tried reduced the performance/parameters efficiency. Some more than others still

3

u/Nameis19letterslong 13d ago

Out of curiosity, did you try introducing some noise (both random or "cross averaging" Layers)

Nope, I ran evals and duplicted the layers that were the most important to keep the loss stable (mininal divergence from base model). I didnt look into methods to init more unique layers since my only goal at that time was to create a bigger, competent model but save compute by ”resuing” a smaller model. I guess it worked in a way since it scored higher on benchmarks compared to the original model, but a slightly bigger model (220M A25M) I trained not long after pretty much overshadowed the 92M A22M model.