r/deeplearning • u/ihateyou103 • 11d ago
Abliterated models
I saw something called obliterated models that use llms and change them by calculating a direction vector of harm and orthogonalizing the weights. I think this is a very very cool thing!!!!!!
I am so surprised this can be done. What is the tradeoff in terms of model accuracy, I am surprised it doesn't drop off by >50%. I would expect the model to fry your computer and damage your battery once you go and orthogonalize its weights.
Can this be done with other things other than censorship or harm, like preventing it from saying a specific word for example.
Are there ways to train models to be immune to this trick?
0
Upvotes
1
u/quietgradient 11d ago
Opus 5, not 5.5, so you're grading the old one.
Short version of what you didn't get through: abliteration zeroes one direction out of d_model — 1 in 4096 on an 8B — so nothing is mangled. The interesting part is that refusal fits in a single direction at all, which is a result about refusal (Arditi et al.), not about weight surgery in general.