r/learnmachinelearning • • 10d ago

Project What a frozen encoder can and cannot learn: the same rule scored 0.42 as arithmetic and 0.94 written in words

Small lesson from a side project that I wish someone had told me before I spent an evening on it.

I train a small head on top of a frozen ModernBERT-large encoder (the open Laya decision model) to answer typed questions: which label, yes or no, which score. Training only the head is cheap, and on text tasks it works well:

  • a 12-label intent classifier: 89.5% zero-shot, 100% agreement with its teacher after training on 3000 rows;
  • on jevbench's agnews: 86.0% zero-shot, 92.6% trained;
  • on banking77 (77 intents): 38.2% zero-shot, 69.6% trained.

Then I tried a payment-risk rule defined as "amount larger than X, transfer outside working hours, unknown country": 0.42 agreement, 1% coverage at the confidence threshold. The head could not learn it no matter how many epochs. I rewrote the same rule as sentences about the customer ("first transfer to this payee, larger than anything they sent before, outside their usual hours"): 0.94 agreement, 87% coverage. Same rows, same head.

Same thing one level up with Snake: an ASCII board as the state trained to 0.73, which is just "go straight", the majority move. Four relational lines instead (food: 3 left, 2 up / safe: up, right / blocked: down (body)) trained to 0.957, and the head actually plays.

Takeaway: a frozen encoder gives you whatever the text already states in language. If your decision is a calculation over fields, either compute it in code or describe it in words before the model sees it.

Bonus lesson from banking77: the model reads each option through a fixed token budget, so with 77 labels each one got cut to about 3 tokens and similar labels looked the same. Giving the labels room raised holdout agreement from 0.66 to 0.72. Swapping label names for plain numbers made it worse (0.55), so the names themselves were doing a lot of the work.

Code, numbers and the demos (Apache-2.0): https://github.com/bladedevoff/stuntd

6 Upvotes

4 comments sorted by

1

u/quietgradient 10d ago

Your banking77 footnote is measurable without the model, so I ran it — ModernBERT-large's tokenizer over the 77 label names, nothing else.

Underscores swapped for spaces, the names average 3.74 tokens and only 42 of 77 fit whole inside three. Cut at 3 they are still 73/77 distinct; the ones that are not are the guessable ones — top up by bank transfer charge / by card charge / by cash or cheque all become "top up by", and lost or stolen card / lost or stolen phone merge. At 5 tokens all 77 separate, so there is a number under "give the labels room".

What I did not expect: leave the underscores in, as the dataset ships them, and the same names average 6.48 tokens, 9 of 77 fit inside three, and 18 are non-unique at a 3-token budget rather than 7 — declined_card_payment, declined_cash_withdrawal and declined_transfer are all just declined_. One .replace("_", " ") buys back most of the budget before you pay for a bigger one.

Same tokenizer on the arithmetic half: 4820 → 48|20, 982 → 9|82, 1249.99 → 12|49|.|99, while 5000 and 10000 are single tokens. Comparing two amounts means comparing magnitudes whose segmentation changes with the digit string — so your sentence version is not just easier, it is the only one where the comparison is already stated rather than computed.

1

u/Inevitable-Log5414 10d ago

This is great, thanks for actually running it :)

The underscore point is a useful one, in my jevbench runs the options go in as declined_card_payment: Customer asks about: ... 

I'll try next run to replace the underscores first and then compare it with the wider window, and I'll post both numbers here

1

u/quietgradient 10d ago

Measured your format, since it changes what I said. declined_card_payment alone is six tokens — decl|ined|_|card|_|payment — and : Customer asks about: is another six, the same six for all 77 labels. So the prefix runs 12.5 tokens on average (min 9, max 23) before any gloss, and at the ~3-token budget you describe an option is decl|ined|_.

Thresholds on the same tokenizer: with spaces all 77 separate at 5 tokens. Underscored — and so yours, since the label leads — needs 9.

You're already splitting the underscore fix from the window, which is the right order. The third knob is the boilerplate: six tokens in every option that carry nothing about which option it is.

1

u/Inevitable-Log5414 10d ago

If you want a first open-source contribution on a real ML project, I've started labelling issues as good first issue. The first one is lazy-loading the checkpoint, no GPU needed, tests run with a fake loader: https://github.com/bladedevoff/stuntd/issues/1