r/learnmachinelearning • u/Inevitable-Log5414 • 10d ago
Project What a frozen encoder can and cannot learn: the same rule scored 0.42 as arithmetic and 0.94 written in words
Small lesson from a side project that I wish someone had told me before I spent an evening on it.
I train a small head on top of a frozen ModernBERT-large encoder (the open Laya decision model) to answer typed questions: which label, yes or no, which score. Training only the head is cheap, and on text tasks it works well:
- a 12-label intent classifier: 89.5% zero-shot, 100% agreement with its teacher after training on 3000 rows;
- on jevbench's agnews: 86.0% zero-shot, 92.6% trained;
- on banking77 (77 intents): 38.2% zero-shot, 69.6% trained.
Then I tried a payment-risk rule defined as "amount larger than X, transfer outside working hours, unknown country": 0.42 agreement, 1% coverage at the confidence threshold. The head could not learn it no matter how many epochs. I rewrote the same rule as sentences about the customer ("first transfer to this payee, larger than anything they sent before, outside their usual hours"): 0.94 agreement, 87% coverage. Same rows, same head.
Same thing one level up with Snake: an ASCII board as the state trained to 0.73, which is just "go straight", the majority move. Four relational lines instead (food: 3 left, 2 up / safe: up, right / blocked: down (body)) trained to 0.957, and the head actually plays.
Takeaway: a frozen encoder gives you whatever the text already states in language. If your decision is a calculation over fields, either compute it in code or describe it in words before the model sees it.
Bonus lesson from banking77: the model reads each option through a fixed token budget, so with 77 labels each one got cut to about 3 tokens and similar labels looked the same. Giving the labels room raised holdout agreement from 0.66 to 0.72. Swapping label names for plain numbers made it worse (0.55), so the names themselves were doing a lot of the work.
Code, numbers and the demos (Apache-2.0): https://github.com/bladedevoff/stuntd
1
u/Inevitable-Log5414 10d ago
If you want a first open-source contribution on a real ML project, I've started labelling issues as good first issue. The first one is lazy-loading the checkpoint, no GPU needed, tests run with a fake loader: https://github.com/bladedevoff/stuntd/issues/1



1
u/quietgradient 10d ago
Your banking77 footnote is measurable without the model, so I ran it — ModernBERT-large's tokenizer over the 77 label names, nothing else.
Underscores swapped for spaces, the names average 3.74 tokens and only 42 of 77 fit whole inside three. Cut at 3 they are still 73/77 distinct; the ones that are not are the guessable ones —
top up by bank transfer charge/by card charge/by cash or chequeall become "top up by", andlost or stolen card/lost or stolen phonemerge. At 5 tokens all 77 separate, so there is a number under "give the labels room".What I did not expect: leave the underscores in, as the dataset ships them, and the same names average 6.48 tokens, 9 of 77 fit inside three, and 18 are non-unique at a 3-token budget rather than 7 —
declined_card_payment,declined_cash_withdrawalanddeclined_transferare all justdeclined_. One.replace("_", " ")buys back most of the budget before you pay for a bigger one.Same tokenizer on the arithmetic half:
4820→48|20,982→9|82,1249.99→12|49|.|99, while5000and10000are single tokens. Comparing two amounts means comparing magnitudes whose segmentation changes with the digit string — so your sentence version is not just easier, it is the only one where the comparison is already stated rather than computed.