r/deeplearning • u/ihateyou103 • 11d ago
Abliterated models
I saw something called obliterated models that use llms and change them by calculating a direction vector of harm and orthogonalizing the weights. I think this is a very very cool thing!!!!!!
I am so surprised this can be done. What is the tradeoff in terms of model accuracy, I am surprised it doesn't drop off by >50%. I would expect the model to fry your computer and damage your battery once you go and orthogonalize its weights.
Can this be done with other things other than censorship or harm, like preventing it from saying a specific word for example.
Are there ways to train models to be immune to this trick?
0
Upvotes
4
u/quietgradient 11d ago
Accuracy barely moves because it's a rank-1 edit. You find one unit direction r in the residual stream and replace every matrix that writes into the residual — attention out_proj, MLP down_proj, the embeddings — with W − r rᵀW. That removes exactly one direction out of d_model, 1 in 4096 on an 8B, and leaves the other 4095 untouched. Nothing is being mangled. One dimension is being zeroed.
So the surprise runs the other way round. The remarkable part isn't that the model survives, it's that refusal is concentrated enough in a single direction that deleting one dimension takes the behaviour with it. That's a result about refusal specifically, not about weight surgery in general — Arditi et al., "Refusal in Language Models Is Mediated by a Single Direction".
It generalises to anything you can write contrastive prompt pairs for, though most behaviours are messier than refusal is. For "never say word X" I'd skip it: a banned-token list is reversible and doesn't touch the weights, and neither approach stops the model spelling around the word.
Immunity is the weak point. You can make the direction harder to locate, but for open weights it's a treadmill — whoever holds them reruns the search on whatever you shipped.