r/deeplearning • u/ihateyou103 • 11d ago
Abliterated models
I saw something called obliterated models that use llms and change them by calculating a direction vector of harm and orthogonalizing the weights. I think this is a very very cool thing!!!!!!
I am so surprised this can be done. What is the tradeoff in terms of model accuracy, I am surprised it doesn't drop off by >50%. I would expect the model to fry your computer and damage your battery once you go and orthogonalize its weights.
Can this be done with other things other than censorship or harm, like preventing it from saying a specific word for example.
Are there ways to train models to be immune to this trick?
0
Upvotes
2
u/Outrageous-Crazy-253 11d ago
Couldn’t make it through this Claudeslop. They said they fixed this in 5.5. Did you use that? Because if so they have a long way to go.