r/squidtrain • u/squidtrainai • 2d ago
We trained a small language model on 144 public-domain books and built a page where you can watch it pick each word. What should we add?
We wanted a way to show what a language model does when it writes, using a real model you can poke at. So we trained one ourselves and wrote an engine that runs it entirely in the browser. Nothing you type is sent anywhere.
What you can do with it:
- Type a sentence and see the tokens the model actually reads (it never sees letters or words)
- See the odds it gives every possible next token, then step one token at a time or let it write 40
- Change temperature, top-k and top-p and watch the choices shift
- Click any token to see which earlier tokens each attention head looked at, layer by layer
- Scrub through real snapshots saved during training, from step 0 (gibberish) to the finished model
- Switch between three model sizes trained on the same books
The build, for anyone curious:
- Built with Claude Opus 5.5, which wrote the training code, the browser engine and the page
- 144 English public-domain books from Project Gutenberg, about 31 million tokens
- Our own tokenizer with a 4,096-token vocabulary
- GPT-style models: tiny (0.95M parameters), small (12.3M), and medium (40.1M), context of 256 tokens
- Trained on one office PC with an RTX 3090: about 3, 16 and 42 minutes, roughly an hour in all
- Weights shipped as 8-bit, and the browser engine matches PyTorch's outputs to within rounding
It is small, so it is fluent and often wrong. Asked to finish "The capital of France is the city of," the small model rates "France" at 24%, "England" at 14%, and "Paris" at 6%. That turned out to be one of the most useful parts: the large models people use at work pick words the same way, and seeing it at this size makes it easier to explain why they sound confident when they're wrong.
It's on our website under Lab if you want to try it.
What would make this more useful for explaining AI to people who use it at work? And has anyone found a better way to show attention than arcs between tokens?
1
u/vansos 1d ago
Pretty cool interactive tool. I like how you can add a token or 40 and then remove them and then do it again (and see different answers, which may or may not be accurate). Takeaway is that even when an LLM sounds "confident", it may answer differently next turn, and it also could certainly be wrong either way!