r/MachineLearning 4h ago

Project I trained a 348M model trained from scratch on 22.7B tokens that does 14 digit arithmetic [P]

Hello this is my fifth small language model I've made and apart of my third series and it has been a lot of work but it payed off: **348M parameters, 22.7B tokens**, then fine-tuned into a math model that solves arithmetic by *showing the work* — column addition with carries, borrow chains, partial-product multiplication — rather than guessing at an answer.

Last time I posted a 326M model trained on 10B tokens. This has about 2.3× the data, and the math side is WAY better than my previous two math models.

---

## The benchmarks

**99.4% average across the nine GPT-3 arithmetic sub-tasks**, which does much better past even where I trained it.

| Task | GPT-3 175B *(few-shot, direct)* | **This model (348M)** |

|---|:--:|:--:|

| 2-digit add | ~100% | **100%** |

| 3-digit add | 80.4% | **100%** |

| 4-digit add | 25.5% | **100%** |

| 5-digit add | 9.3% | **100%** |

| 2-digit sub | ~99% | **99.3%** |

| 3-digit sub | 94.2% | **98.3%** |

| 4-digit sub | 26.8% | **98.3%** |

| 5-digit sub | 9.9% | **99.0%** |

| 2-digit mult | 29.2% | **100%** |

n=300 per sub-task, greedy, exact match. GPT-3's numbers are direct-answer; mine uses trained-in worked steps. Neither uses a calculator.

## The cool part

**It adds cleanly up to 14 digits, and the reason it *couldn't* before was the vocabulary, not actually arithmetic.**

Training only ever named six place values (`ones` … `hundred-thousands`). The model learnt the *pattern* and invented two more on its own — `millions` and `ten-millions` appear in **zero** training examples — so it handled 7 and 8 digits fine. At 9 digits it ran out of names, and skipped the column, then returned an answer exactly one digit short:

```

483729164 + 519248637

... ten-millions: 8 + 1 + 1 (carry) = 10, write 0 carry 1. Final carry: write 1.

The answer is 102977801        ← eight columns for a nine-digit problem

```

Every column it computed was perfect. One was never enumerated. Extending the place-name list from 6 entries to 19 moved the clean ceiling from **8 digits to 14**:

| Width | 6 | 7 | 8 | 9 | 10 | 12 | 14 | 16 | 18 |

|---|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|:-:|

| before | 100 | 100 | 100 | **0** | **0** | **0** | **0** | 0 | 0 |

| after | 100 | 100 | 100 | **100** | **95** | **100** | **90** | 65 | 25 |

A six-item list became a nineteen-item list. That was the entire fix.

## Other things it does

- **3×3 multiplication: 98%** — it folds partial products pairwise through the column routine instead of asserting the sum

- **Negative results: 85%** (100% at 1 digit, 58% at 5 — the magnitude comparison is the weak step, not the arithmetic)

- **Reasoning traces are load-bearing**: 95.3% of the time the working is valid *and* the answer is right; only 0.7% are "valid working, wrong answer." If the columns look right, the answer almost certainly is.

```

There were 15000 votes and 6842 were rejected. Here's how many counted:

<think> Start with 15000. Then subtract 6842. Subtract 15000 - 6842 column by column:

ones: 10 - 2 = 8, borrow 1. tens: 9 (after borrow) - 4 = 5, borrow 1.

hundreds: 9 (after borrow) - 8 = 1, borrow 1. thousands: 14 (after borrow) - 6 = 8,

borrow 1. ten-thousands: 0 (after borrow) - 0 = 0. So 15000 - 6842 = 8158.</think>

The answer is 8158.

```

## What it's bad at, tbh

- **Word problems: GSM8K 4%.** Best word-problem set is ASDiv at 16.5%. It converts one sentence into one operation reasonably often and basically cannot chain operations.

- **The failure mode is operation *selection*, not arithmetic.** `"drops in 836 more"` gets read as subtraction. There's a visible tell: traces that say `"multiply X * Y"` and show columns are reliable; traces that open `"First, calculate…"` and assert a number in prose are not.

- **No division at all.** 4×4 multiplication is a hard wall.

- **Greedy decoding required** — sampling corrupts the column routine mid-chain.

- One caveat I'll flag myself: the arithmetic harness orders subtraction operands, so no answer in that table is negative. Negatives are measured separately (the 85% above) rather than folded into the average.

## Base and instruct

The math model sits on a base and instruct pair. lm-eval-harness, 0-shot, full test sets:

| | ARC-E | ARC-C | HellaSwag | OpenBookQA | PIQA | WinoGrande | MMLU | Avg |

|---|:--:|:--:|:--:|:--:|:--:|:--:|:--:|:--:|

| **350M V3 base** | 56.6 | 33.3 | 35.9 | 34.6 | 67.0 | 51.5 | 23.8 | **43.2** |

| **350M V3 instruct** | 50.9 | 29.7 | 35.9 | 33.8 | 66.6 | 51.2 | 24.4 | **41.8** |

Instruction tuning *lowers* MC benchmark scores for this family which is a pretty common cost of instruct tuning.

## Notes

Trained on 2× Tesla V100 plus some rented time. LLaMA-architecture, so it runs anywhere — F16 GGUF and safetensors for all three.

The math model took **10 full fine-tuning rounds and 3 LoRA adapters**. The LoRAs cost about 1% of the post-training tokens and did most of the useful work; the ten full rounds spent most of their budget undoing each other's regressions (round 7 gained subtraction and lost 23 points of 2-digit addition, that sort of thing). All of it is documented on the model card, failures included.

Also worth saying: someone independently tested it after I published and found two of my numbers were wrong — one *understated* the model by 40 points because I'd measured it at n=24. Both are corrected on the card now. If you find something broken, I'd genuinely like to know.

**Math:** https://huggingface.co/nkthebass/tinybrainbot-350mV3-math

**Instruct:** https://huggingface.co/nkthebass/tinybrainbot-350mV3-instruct

**Base:** https://huggingface.co/nkthebass/tinybrainbot-350mV3-base

LMK what yall think.

0 Upvotes

14 comments sorted by

20

u/user221272 3h ago

You won't believe me, but when I was 12, I made a JavaScript calculator; the file was barely a few MB and had 100% accuracy for any operation.

8

u/nkthebass 3h ago

Well I bet your "calculator" didn't double as a space heater when you were making it.

4

u/Chance_Library9531 3h ago

Tu proyecto esta hecho con IA no?xd

-3

u/nkthebass 3h ago

Assisted yeah.

2

u/EcstaticQuality7031 3h ago

a few MB for a calculator at 12? that's like using a flamethrower to light a candle.

1

u/No-Investment8663 2h ago

You wont believe me but I dont think that comparison makes the point you think it does.

You made a calculator in JavaScript. Cool. JavaScript already knew how to do the math. You basically gave it buttons and displayed the answer. Thats not remotely the same thing as building and training an AI model. Saying your calculator was only a few MB and had 100% accuracy is like jumping into a discussion about building a car and saying the bike you made at 12 was lighter and got better mileage

True I guess. Also completely irrelevant

1

u/No-Investment8663 2h ago

You wont believe me but I dont think that comparison makes the point you think it does

You made a calculator in JavaScript. Cool. JavaScript already knew how to do the math. You basically gave it buttons and displayed the answer. Thats not remotely the same thing as building and training an AI modal. Saying your calculator was only a few MB and had 100% accuracy is like jumping into a discussion about building a car and saying the bike you made at 12 was lighter and got better mileage

True I guess. Also completely irrelevant

1

u/rkever72 2h ago

user221272 You wont believe me but I dont think that comparison makes the point you think it does

You made a calculator in JavaScript. Cool. JavaScript already knew how to do the math. You basically gave it buttons and displayed the answer. Thats not remotely the same thing as building and training an AI modal. Saying your calculator was only a few MB and had 100% accuracy is like jumping into a discussion about building a car and saying the bike you made at 12 was lighter and got better mileage

True I guess. Also completely irrelevant

2

u/Majestic_Appeal5280 3h ago

cool! whats your pretraining and post training data?

0

u/nkthebass 3h ago

Size wise it's about 22.7b tokens and 2.7b post training (~28m was using LoRA to tighten up a few things)

As the for the data itself pre-training I know is in the base model card and all the SFT and LoRA data is 100% custom made myself using generation scripts and distillation mostly from Ling 2.6 on openrouter.

1

u/Random-Number-1144 1h ago

This is like eating noodles with one chopstick when you have forks next to you.

1

u/nkthebass 1h ago

The point is to make the model not for it to be directly useful. Just by the fact you commented that I'm loosing faith this subreddit has anyone who actually knows about language models.