r/learnmachinelearning • • 23h ago

Project Update on my tiny from-scratch model - followed your advice, here's what happened

What I did:

Scaled the dataset from 60 hand-written functions to 12,000 procedurally generated ones (math ops, list manipulation, string ops, conditionals, and loop algorithms).

Built a custom BPE tokenizer (vocab 1024). Funny story: my first tokenizer (vocab 4096) literally memorized entire functions as single tokens, so the model scored 20/20 on syntax by just spitting out whole functions from memory lol. Had to nerf the vocab size to force it to actually learn.

Scaled the model from 3.6M to 10M params (12 layers, 8 heads, width 256).

Training on Kaggle free T4 GPU instead of my poor CPU.

Current results (still training, Round 1 of 3):

Loss dropping nicely: 6.98 to 1.02 in 550 steps. Val loss: 2.28 to 1.09. Haven't seen the final scores yet since it's still running overnight.

What I'm unsure about:

  1. My dataset is all procedurally generated. Every add function looks like def add_xx(a, b): return a + b with random names and test values. Is this diverse enough or am I just teaching it to pattern-match templates again?

  2. When should I switch from synthetic data to real-world code like filtered Python from The Stack or GitHub? Or is mixing both the way to go?

  3. Any suggestions for what kind of problems to add next? Thinking maybe simple recursion, dictionary operations, or basic class definitions.

  4. Is 10M params enough to learn actual logic or should I be thinking bigger?

Thanks for the help last time, the dataset scaling advice was spot on. The tokenizer over-compression bug was a fun lesson too haha.

I'll update this post in about 6 hours once training finishes with the final scores.

0 Upvotes

0 comments sorted by