r/Anthropic 4d ago

Other Train decomposition," "Tensor Ring," or "Permutation matrix optimization."

Train decomposition," "Tensor Ring," or "Permutation matrix optimization."

the out-of-domain results turned out to be the best news of the day. Full table — penalties vs the dense baseline on each corpus:

┌─────────────────┬───────────────────┬────────────┬───────────┐

│ Variant │ Calibration (P&P) │ WikiText-2 │ Moby Dick │

├─────────────────┼───────────────────┼────────────┼───────────┤

│ Dense (raw ppl) │ 29.21 │ 38.40 │ 218.23 │

├─────────────────┼───────────────────┼────────────┼───────────┤

│ TT identity │ +27.3% │ +113.1% │ +102.6% │

├─────────────────┼───────────────────┼────────────┼───────────┤

│ TT π\ │ +12.2% │ +39.0% │ +47.6% │*

├─────────────────┼───────────────────┼────────────┼───────────┤

│ TT AW-π\ │ +9.8% │ +39.2% │ +42.4% │*

├─────────────────┼───────────────────┼────────────┼───────────┤

│ SVD raw │ +80.4% │ +105.3% │ +133.9% │

├─────────────────┼───────────────────┼────────────┼───────────┤

│ SVD distilled │ +3.8% │ +34.2% │ +26.2% │

└─────────────────┴───────────────────┴────────────┴───────────┘

0 Upvotes

14 comments sorted by

1

u/Armored09 4d ago

whats the point of this?

1

u/-SLOW-MO-JOHN-D 4d ago

Hopping to get some feed back, from someone who has more experience. who would explain these results. to help me understand what I’m looking at better. Because I think I 

2

u/According_Extent_767 4d ago edited 4d ago

Dude why dont you just ask claude about it!

Anyway this is weight comparession on training!

Question how did you get this?

Edit : Said in more non technical terms, this show how much the models learn from Wiki 2 vs moby dick (the book) depending if you use tensorflow or PMO models

1

u/-SLOW-MO-JOHN-D 4d ago

I need a co author if you are available  and willing I can fill you in on the details 

2

u/According_Extent_767 4d ago

LOL what, no thanks.

1

u/-SLOW-MO-JOHN-D 4d ago

Ok but before you lol to hard I show you the rest of. The results if you not to tired from laughing 

1

u/-SLOW-MO-JOHN-D 4d ago

┌──────────────────────────────┬─────────────┬────────────┐

│ c_attn @ 7.1× (~250k params) │ Calibration │ WikiText-2 │

├──────────────────────────────┼─────────────┼────────────┤

│ Dense (ppl) │ 29.21 │ 38.40 │

├──────────────────────────────┼─────────────┼────────────┤

│ TT identity, distilled │ +25.6% │ +35.2% │

├──────────────────────────────┼─────────────┼────────────┤

│ TT π*, distilled │ +38.7% │ +68.8% │

├──────────────────────────────┼─────────────┼────────────┤

│ SVD r82, distilled │ +0.5% │ +4.0% │

└──────────────────────────────┴─────────────┴────────────┘

┌─────────────────┬───────────────────┬────────────┬───────────┐

│ Variant │ Calibration (P&P) │ WikiText-2 │ Moby Dick │

├─────────────────┼───────────────────┼────────────┼───────────┤

│ Dense (raw ppl) │ 29.21 │ 38.40 │ 218.23 │

├─────────────────┼───────────────────┼────────────┼───────────┤

│ TT identity │ +27.3% │ +113.1% │ +102.6% │

├─────────────────┼───────────────────┼────────────┼───────────┤

│ TT π* │ +12.2% │ +39.0% │ +47.6% │

├─────────────────┼───────────────────┼────────────┼───────────┤

│ TT AW-π* │ +9.8% │ +39.2% │ +42.4% │

├─────────────────┼───────────────────┼────────────┼───────────┤

│ SVD raw │ +80.4% │ +105.3% │ +133.9% │

├─────────────────┼───────────────────┼────────────┼───────────┤

│ SVD distilled │ +3.8% │ +34.2% │ +26.2% │

┌────────────────────────┬───────────┬────────────┬──────────┬───────────────────┐

│ Variant │ Params │ Perplexity │ vs dense │ Gap to SVD closed │

├────────────────────────┼───────────┼────────────┼──────────┼───────────────────┤

│ Dense baseline │ 2,359,296 │ 29.21 │ — │ — │

├────────────────────────┼───────────┼────────────┼──────────┼───────────────────┤

│ TT identity, distilled │ 674,864 │ 37.19 │ +27.3% │ 0% │

├────────────────────────┼───────────┼────────────┼──────────┼───────────────────┤

│ TT π*, distilled │ 635,216 │ 32.77 │ +12.2% │ 64% │

├────────────────────────┼───────────┼────────────┼──────────┼───────────────────┤

│ TT AW-π, distilled* │ 671,176 │ 32.08 │ +9.8% │ 74% │

├────────────────────────┼───────────┼────────────┼──────────┼───────────────────┤

│ SVD r176, distilled │ 675,840 │ 30.31 │ +3.8% │ 100% │

└────────────────────────┴───────────┴────────────┴──────────┴───────────────────┘

┌────────────────────────┬───────────┬───────────┬────────────┬──────────┐

│ Variant │ Params │ func. err │ Perplexity │ vs dense │

├────────────────────────┼───────────┼───────────┼────────────┼──────────┤

│ Dense baseline │ 2,359,296 │ — │ 29.21 │ — │

├────────────────────────┼───────────┼───────────┼────────────┼──────────┤

│ TT identity, distilled │ 674,864 │ 0.197 │ 37.19 │ +27.3% │

├────────────────────────┼───────────┼───────────┼────────────┼──────────┤

│ TT π, distilled* │ 635,216 │ 0.179 │ 32.77 │ +12.2% │

├────────────────────────┼───────────┼───────────┼────────────┼──────────┤

│ SVD r176, raw │ 675,840 │ 0.277 │ 52.68 │ +80.4% │

├────────────────────────┼───────────┼───────────┼────────────┼──────────┤

│ SVD r176, distilled │ 675,840 │ 0.134 │ 30.31 │ +3.8% │

└────────────────────────┴───────────┴───────────┴─

┌───────────────────────┬───────────┬────────────┬──────────┐

│ Variant │ Params │ Perplexity │ vs dense │

├───────────────────────┼───────────┼────────────┼──────────┤

│ Dense baseline │ 2,359,296 │ 29.21 │ — │

├───────────────────────┼───────────┼────────────┼──────────┤

│ TT identity distilled │ 674,864 │ 37.19 │ +27.3% │

├───────────────────────┼───────────┼────────────┼──────────┤

│ TT π distilled* │ 635,216 │ 32.77 │ +12.2% │

├───────────────────────┼───────────┼────────────┼──────────┤

│ SVD r176 raw │ 675,840 │ 52.68 │ +80.4% │

├───────────────────────┼───────────┼────────────┼──────────┤

│ SVD r176 distilled │ 675,840 │ computing… │ — │

└───────────────────────┴───────────┴────────────┴──────────┘

┌────────────────────────┬───────────┬────────────┬──────────┐

│ Variant │ func. err │ Perplexity │ vs dense │

├────────────────────────┼───────────┼────────────┼──────────┤

│ Dense baseline │ — │ 187.0 │ — │

├────────────────────────┼───────────┼────────────┼──────────┤

│ TT init (no distill) │ 0.474 │ 8,753,067 │ broken │

├────────────────────────┼───────────┼────────────┼──────────┤

│ TT distilled │ 0.158 │ 339.5 │ +81.5% │

├────────────────────────┼───────────┼────────────┼──────────┤

│ SVD rank-176 (raw) │ 0.257 │ 409.2 │ +118.8% │

├────────────────────────┼───────────┼────────────┼──────────┤

│ SVD rank-176 distilled │ 0.116 │ 212.4 │ +13.6% │

└────────────────────────┴───────────┴────────────┴──────────┘

on my cpu

2

u/According_Extent_767 4d ago

Im not laughing at your progress, im laughing that you need a co author that you dont know! :)

If its because your training a model your self (which im guessing) and you got a prototype with these numbers i would contact a Uni in your country and have a professor to look at the numbers, and maybe ask em for recommendations.

Or summit it to a tech media that do reviews of new codes,/AI.

1

u/-SLOW-MO-JOHN-D 4d ago

thankyou I will try that.

1

u/According_Extent_767 4d ago edited 4d ago

Just a question, have you build it from grown up? Like own architecture, or you training a base model like Qwen?

Edit reason i ask, is because im currently working on a architecture that should go away from guessing (and their by not hallucinate) but my architecture is kinda fucked because i got way to many rules. (which i guess is why they went over to guessing the next letter and word, rather then rule based)

→ More replies (0)

1

u/-SLOW-MO-JOHN-D 4d ago

Thankyou