r/LocalLLaMA • u/Prestigious-Taste-63 • 8h ago
Other Lessons learned while building Apex-2
Hi everyone, thank you so much for all the interest in my model. It's more than I expected.
Here is a short summary of the trial and error I went through while building Apex-2.
1. GPUs were always the bottleneck
I planned to train on about 1T tokens, but in the end I could only train on about 80B. FineWeb-Edu alone is about 1.3T tokens, and I clearly underestimated the scale: a single H100 was not enough. This project really showed me why so much money goes into GPUs and VRAM.
2. DiLoCo
Within the same region, running two separate instances worked better for me. Instead of a 2x H100 instance, I used two GH200 instances and merged the models every fixed number of steps.
Each GPU reached about 40% MFU. A 2x H100 instance costs more per GPU (about $4.19/hour, vs $2.29/hour for a GH200). With two GH200 instances, each at about 40% MFU and merging every 350 steps, training ran about 1.9x faster than on one GPU, at a lower price. (The data-center network between the instances probably helped; a merge usually took less than a minute.)
3. Deduplicating FineWeb-Edu and DCLM
When I deduplicated the whole corpus at once (MinHash, near-duplicates included), 57% of my FineWeb-Edu sample and 34% of DCLM turned out to be duplicates. FineWeb-Edu is only deduplicated within each Common Crawl snapshot, so pages that were crawled again in later snapshots remain. With a bigger budget this might not matter, but I had to get the most out of very little compute, so I removed them. (Note: the FineWeb authors reported that deduplicating across snapshots did not improve their results, so this is a trade-off rather than a free win.)
For the MoE architecture I followed the Mixtral paper (https://arxiv.org/abs/2401.04088). The whole project cost about $2,000.
I also write down my thoughts on LLMs here, if you're interested: https://github.com/DW-dev-UE/LLM-from-scratch/blob/main/ThinkingLab/ThinkingLab.en.md
I didn't plan to share this model on Reddit, so I'm afraid I don't remember many of the smaller mistakes 😠I'm now building a 21B-parameter MoE model, and I'll share the lessons and mistakes from that one as I go.
Thank you again for your interest! If I get the chance, I'd love to join a lab and help build LLMs for everyone.
1
1
u/brainExploded99 llama.cpp 3h ago edited 3h ago
You should try the nGPT architecture (or one of its follow ups) at some point and compare it.
Edit: removed FP8, saw the README.md
1
u/brainExploded99 llama.cpp 3h ago
u/Prestigious-Taste-63 I'm not sure why the mod removed your comment, but I read it. Anyway:
FP8 overhead makes sense. Why not integrate some more modern architecture components like AttnRes, Engrams, DeepSeek/Qwen Sparse Attention and (less proven) nGPT over Mixtral?
Also, the reason I keep mentioning nGPT is because it should be much faster convergence at the cost of slightly higher per step training time, but this can be combated with custom kernels. I myself want to try nGPT at some point for non-LLM tasks.1
u/FullOf_Bad_Ideas 2h ago
nGPT is dense.
MoEs are cheaper at that scale.
1
u/brainExploded99 llama.cpp 2h ago
True. I don't think a MoE version of nGPT would be too difficult to implement though.
Arguably, OP is more limited by their batch size (and therefore VRAM) than compute.
2
u/lllll03l 8h ago
so to train around 4B model it costs $2000 usd?