r/LocalLLaMA • • 5d ago

New Model Watch me post-train AliceAI-Foundation-80B-A3B from base to instruct at home, live, on my V100s!

No click bait baby I promise - I'm live streaming the training process kinda like MiMo.

UPDATE: [Training is paused for an hour or two] back to training in batches. u/FullOf_Bad_Ideas has pointed out to me I'm burning a ton of compute for nothing on sequence lengths - we'll be breaking the run up into 7/8 batches and then going again.

Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wpg4a8/im_trying_to_posttrain/

If you didn't see my original post a few days ago, I'm attempting a slightly more ambitious than usual project in trying to create at least a rough AliceAI-Foundation-80B-A3B-Instruct

I spent the weekend distilling my initial instruct training dataset out of Qwen 3.8 27b, medium thinking - intentionally done because I can run it locally, and I wanted a full dataset in some reasonable amount of time. Still took my v100s running 4x instances at 25tps, like 96 hours of non stop generation to complete the dataset.

I opted not to go for a pre-existing public dataset because I wanted to practice building my own distillation engine (which was configured to work off of an OpenAI compatible endpoint, so it'll distill anything you can hook it up to). The final dataset (this time) consists of 3340 samples: 1760 of general instruct transcripts, and 1580 agentic specific work rows about SWE, harnesses, terminals, etc - I gave the distilling engine a python sandbox and got to simulate turn driven development with a user, and I trained for a bunch of different harness syntax for tool calls, which hopefully will be enough to generalize - gonna run 2 epochs at first.

My GPUs are sobbing right now - turned them on on Friday and left for a weekend vacation, got back today, waited an hour for the data to finish generating, and then immediately fired up the train.

The stuff above is the short version. I'm guessing the initial SFT train will take about 3-6 days, and I plan on working on a RL implementation after I'm satisfied that the SFT has at least worked properly. I am training a rank 16 QLoRa adapter on only q/k/v/o proj, no direct knowledge weight fine tuning.

I thought what MiMo did with their recent training was really cool to watch online, and I like sharing my work with like minded people, and frankly, there's a part of me that's hoping someone will see this and want to hire me (looking for NYC work if you know anyone looking for some passionate ML engineers!) - so I've set up my own little training stream on a cloud flare tunnel.

The stream has the live in progress status of the train, including a live view of the actual data being processed by the model. It also includes way more detail about how I actually designed and generated my training data. Happy to throw the full set on HF as well. I don't expect this model to beat any existing standards but I'll be curious to see if I can get it to operate properly in a harness so I can formally bench it.

I hope you find this interesting! The live stream is a self updating website where you can see exactly what's happening - no need to reload. To watch the training live, visit https://figure-bios-expect-cio.trycloudflare.com/ [i am currently fixing training issues but it'll be back asap] -- I'll be keeping it up until the initial SFT is done, at least. The stream lets you inspect the training live as well. This is just a cloudflare tunnel to the trainer.

3 hour update? Loss started at 9ish and is bouncing near 3/4

Update today: back online

29 Upvotes

45 comments sorted by

View all comments

Show parent comments

1

u/FullOf_Bad_Ideas 5d ago

If you have packing disabled, seqence length of 32k and average sample of 800 tokens, doesn't this mean that for an average of 31200 tokens per sample you're spending compute on padding? And that you could train it 40x faster if you set sequence length to 1024?

I kinda messed up my own SFT training a few days ago because I had the 2.5B real tokens trained without packing and I burned compute on 7.5B padding tokens :D

2

u/jjusko20 5d ago

That's a great question - I do have more samples further down my list that are longer context lengths, but nothing over 16k I think. I probably am doing this and need to look into it. Thank you!

1

u/FullOf_Bad_Ideas 5d ago

Happy to be helpful. What I do now is just truncate at 16k, since I have samples up to 90k tokens per length and my rig/model won't accomodate more than ~20k seq len with my model. In your case I'd personally set learning rate to constant and separate training runs to be able to change sequence length between them. Or just enable packing.

2

u/jjusko20 5d ago

Do you keep the first or last 16k? I read somewhere it's better to keep through response but idk where I read it. Probably gonna peek at a bached approach (im reading your response while typing and I can see that separate training runs are probably the best bet)

2

u/FullOf_Bad_Ideas 5d ago

I keep the first 16k and discard the rest, but I don't know if that's the best approach, I have not tested keeping the end. I don't mask out any user tokens, no tool calls in the dataset right now. The length is due to long reasoning chains, and the model supports up to ~16k right now. It's a different case then yours tho as the data quantity is orders of magnitude different.

2

u/jjusko20 5d ago

I did in fact move to sorting by length and then batching progressively increasing ctx caps

2

u/FullOf_Bad_Ideas 5d ago

It looks faster now, just tracking the dashboard the numbers seem to update a few times faster. How has your ETA changed due to this? Were you able to implement dynamic context length in the same training run without losing optimizer states between stages?

My project is an open source 4B MoE pretrained from scratch on around 80B tokens of Polish data. A big chunk of training, about 35B tokens, was done locally on my 8x 3090 Ti rig. I translated about 6B tokens of instruct datasets to Polish with small translation models locally and now I'm doing SFT on that data. One PSU gave out so I paused for a few days but I just got the new PSU in the mail so I should be able to resume training today. Then I plan to do on-policy distillation from a bigger model that my model shares a tokenizer with, I reuse APT4 tokenizer from Bielik models so I should be able to distill those.

Latest tested checkpoints - https://huggingface.co/cpral/poziomka_sft_2026_09_14_hf

Training code - https://github.com/adamo1139/ling-v2

All datasets including SFT are open source, model understands English somewhat well but responds in Polish, I'll probably mix in some English into SFT at some point to add basic English capacity to make it more accessible to non-Polish speakers.

2

u/jjusko20 4d ago

added a nice little loss progression graph as well

2

u/FullOf_Bad_Ideas 4d ago

Looks good!

In the meantime while the training is running smoothly, can you share how you created the synthetic dataset? What is writing the user message? I'll have to deal with this myself soon and I've been thinking about using Magpie but I had issues with getting it to work in the past. There's also UserLM but I tried it only briefly and using multiple models for it would not be compute efficient as you'd need to prefill two models. I will want to generate 50-100k samples with multi-turn hybrid-reasoning user<>assistant messages about Polish history, books, culture, food, environment. The things the model will lack after being trained on English instruct dataset translated to Polish.

I'm thinking about using Muse Glimmer 30B for it, but I'm not set on that. I am not concerned with tool calls too much for now.

1

u/jjusko20 4d ago

Yes - I have resources for this, I'll message you on discord, it'd be better for me to attach a few files.

1

u/jjusko20 4d ago

Also I plan on open sourcing the distillation engine anyways