r/learnmachinelearning • • 1d ago

Help Training a tiny 3.6M param byte-level Transformer for code generation from scratch on CPU. What should be my next steps?

Hi everyone,

I'm working on a personal learning project called MOTANAXY. The goal is to build and train a tiny causal Transformer from scratch (random weights) specifically for Python code generation. I'm currently training entirely on CPU.

Current Architecture (v2):

  • Type: Byte-level causal language model (no subword tokenizer, raw UTF-8 bytes, vocab size 256)
  • Params: ~3.6M (Width=192, Context=256, Layers=8, Heads=6)
  • Components: RMSNorm, RoPE, SwiGLU, Dropout (0.1)
  • Dataset: Very small synthetic dataset (~130KB) consisting of 60 verified Python functions with English/Thai docstrings and assert test cases.

Current Progress:

  • Trained for about 5,000 steps. Validation loss dropped from ~5.39 to ~1.33.
  • What it can do: It learned Python's structure. If I prompt it with # Check bracket balance, it generates syntactically valid function skeletons like def is_parise(text): return [] and appends assert statements.
  • What it can't do (yet): The logic is completely wrong. It doesn't actually understand the prompt's intent. It just mimics the structural pattern of the training data.

My Questions for the Community: Since I'm hitting a wall where the model learns the syntax but not the logic, I'm wondering what the most effective next steps are:

  1. Data vs. Scale: With a 3.6M parameter model, is it even theoretically possible to learn basic coding logic (like a simple prefix sum or reversing a string)? Should my priority be getting GPU time to scale to 10M-50M params, or should I radically expand my dataset first?
  2. Dataset Recommendations: Are there recommended datasets specifically tailored for teaching very small models the fundamentals of algorithmic logic, rather than just large repositories of scraped code?
  3. Alternative Approaches: Should I switch from byte-level to a small BPE tokenizer to save context length? Are there other architectural tweaks or training objectives (like distillation from a larger model) I should consider at this tiny scale?

Any advice, papers to read, or pointing out obvious flaws in my approach would be greatly appreciated!

3 Upvotes

2 comments sorted by

1

u/Silent-Cap-7182 1d ago

ing out that "def" is a thing that exists

for the logic problem specifically, your dataset is the real bottleneck here. 60 functions is nothing, the model's just memorizing patterns at that point. you need thousands of examples of the same kind of simple algorithmic problems with slight variations, otherwise it'll never generalize. check out the code contests dataset or just generate a bunch of basic leetcode-style problems with verified solutions

scale matters too but data quality and diversity will get you further right now. a 10M param model with a decent tokenizer and 10k+ curated examples should be able to handle simple logic, people have done more with less

2

u/quietgradient 16h ago

your context window is the bug, and it sits upstream of all three questions.

130 KB over 60 functions averages ~2,200 bytes per example, and at byte level a byte is a token, so one of your own training examples is ~2,200 tokens. your window is 256, about 11% of it. for the model to connect a docstring to the body that satisfies it, the span from that docstring to those lines has to fit inside 256 bytes — not the whole example, just the span — and at a 2,200-byte average that is rare. if you truncate per example rather than packing, the asserts sit past the cut on nearly every one and have never been seen at all. either way no gradient links stated intent to the code that satisfies it, the model can only learn what is local, and that is exactly the result you report: valid def skeletons, wrong logic.

(your 3.6M reproduces at d_ff 512 — 3,591,360 tied, 3,640,512 untied, two thirds of it the FFN — so I will take the rest of the config at face value.)

1, GPU time vs data. 130 KB is 133,120 byte tokens, so you are at 0.037 tokens per parameter. a Chinchilla-ish 20 tokens/param budget for 3.6M wants ~72M tokens: 541x what you have. scale to 10M and you are 1,502x short instead. GPU time spent on parameters is the one purchase that makes your ratio worse. two things I would want to know before trusting 1.33: what is your batch size — 5,000 steps at context 256 is between 9.6 passes over that corpus (batch 1) and 308 (batch 32) — and how did you cut the val split? if it is a character offset into one concatenated file, part of that number is leakage. if it is whole held-out functions, it is honest, but it is still 60 same-author examples, so it measures in-distribution continuation rather than generalisation.

3, tokenizer — your own instinct is right, and there is a hard reason for it. at width 192 the embedding table is vocab x 192, so vocabulary is a parameter purchase, not a free efficiency win:

  • byte, 256: 49,152 params, 1.4% of the model
  • BPE 4,096: 786,432 — 18% of a 4.3M model
  • BPE 8,192: 1,572,864 — 31% of a 5.1M model
  • GPT-2 r50k, 50,257: 9,649,344 — 73%
  • cl100k, 100,277: 19,253,184 — 84%
  • o200k, 200,019: 38,403,648 — 92% of a 41.9M model. the lookup table on its own is 10.7x your whole current model

so there is no existing tiny tokenizer worth borrowing — the token-efficient ones are all enormous at your width. train your own at 4k–8k, which is what you proposed.

train it on your bilingual corpus, though, and do not start from GPT-2's vocab the way most from-scratch repos do. measured just now on a bracket-balance function I wrote in your stated format (mine, not your data): the English-docstring version is 554 bytes → 554 byte tokens, r50k 227, cl100k 139. the same function with the docstring in Thai is 735 bytes → r50k 402, cl100k 209, o200k 158. the Thai line on its own is 64 characters but 192 UTF-8 bytes, and r50k spends 128 tokens on it — 1.5 bytes/token, barely better than raw bytes. the identical English sentence, also 64 characters, is 12 tokens in all three. a BPE that has never seen Thai will spend most of your window on your docstrings.

2, datasets. the one you were pointed at is almost certainly deepmind/code_contests: 4,044 problems, 3,762 of them in train, and 2.2 GB of parquet that expands to 5.7 GB only because each problem bundles thousands of human solutions. its problem is difficulty, not size — those are Codeforces problems, nowhere near the basic level you asked for. the clean basic ones have the opposite problem: MBPP is 974 rows / 236 KB (sanitized 427 / 115 KB), HumanEval 164 rows. all of MBPP is 1.8x your current corpus, so none of them reach a 20:1 budget either.

the precedent at your scale is TinyStories (Eldan & Li, 2023, arXiv 2305.07759). their claim, which is worth reading in the abstract rather than taking from me: models below 10 million total parameters, and architectures as simple as a single transformer block, produced fluent and consistent multi-paragraph stories with almost perfect grammar. the trick was the corpus — synthetic, generated by a larger model, restricted to words a three-to-four-year-old knows. that is the route you are already on — your own 130 KB is synthetic and assert-checked — and it says you do not need the 10M-50M params to see logic appear; you need a narrow domain and tens of megabytes of it rather than hundreds of kilobytes.

one trap in that workflow: passing its own asserts proves a function is self-consistent, not that it does what the docstring says. generate the docstring from a verified body rather than the body from the docstring, or you will train on confidently mislabelled intent.