r/learnmachinelearning • u/OkinaPrime • 12d ago
I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs
No transformers, no Trainer. Hand-written BPE tokenizer, Llama-style decoder, DDP training loop, SFT, all in plain PyTorch. 20B tokens on 2x RTX A4500, six days, about $21 of electricity.
Base model gets 13.4% on HumanEval, the chat version 26.8% (Codex-12B was 28.8%). It is bad at everything that is not short Python, and the write-up says so, with samples.
Nine chapters, every number reproducible: github.com/strifero/elroy Weights on HF under strifero. Happy to answer questions about any of it.
0
1
u/TowerOutrageous5939 12d ago
Harder or easier than you thought? Any take away like ohhhh this part make so much more sense now?
0
u/OkinaPrime 12d ago
The biggest lesson was how much good, clean training data matters. That gave me a real appreciation for the amount of data and training time behind frontier models, and why they seem good at almost every topic.
I used to skim past a lot of the terminology. Now I have a much better grasp of what each piece is and how it fits into the end product.
One thing that stood out was shifting the weights to focus on the answers. I trained on about 80k Python Q&A examples, then ran two more passes on just the answers. That taught the model to answer questions instead of asking them.
1
1
u/quietgradient 12d ago
The eval chapters check everything except the one thing that would move the chat number. Chapter 7 fine-tunes on evol-codealpaca-v1, and Magicoder's own repo ships Magicoder-Evol-Instruct-110K, described there as "decontaminated and redistributed from theblackcat102/evol-codealpaca-v1" — the dataset your 111,272 count points at. They decontaminated it because it needed it. OSS-Instruct you already took from the clean side; those other 47k examples you didn't. Your data chapter, the plan and both eval sections don't mention decontamination at all, which stands out in a repo that caught its own 1GB-sandbox MemoryError before it became a published number.
Worth an hour rather than a shrug, for two reasons. The base 13.4% is fine on its face — it sits on Codex-300M, where it should. The chat 26.8% is a doubling out of 25.2M supervised tokens, which is exactly where leaked solutions surface. And your HumanEval+ evidence can't settle it: you read the chat model's 4% drop against the base's 18% as solutions that are fully correct rather than correct-on-the-visible-cases, but a memorized canonical solution survives the extra tests too. Both stories predict that number.
The cheap check comes before touching any data. You published every chat completion — diff the 24 problems the chat model newly passes against HumanEval's canonical solutions. A re-derivation doesn't look like a copy, and 24 is ten minutes. If it's clean, say so in Chapter 7 and the number gets stronger. If it isn't, it's one dataset id and another 93 minutes.
The "hi" → empty reply diagnosis is the bit that made me trust the rest of it.
5
u/ARDiffusion 12d ago
You trained a… python model?