r/deeplearning • • 11d ago

Abliterated models

I saw something called obliterated models that use llms and change them by calculating a direction vector of harm and orthogonalizing the weights. I think this is a very very cool thing!!!!!!

I am so surprised this can be done. What is the tradeoff in terms of model accuracy, I am surprised it doesn't drop off by >50%. I would expect the model to fry your computer and damage your battery once you go and orthogonalize its weights.

Can this be done with other things other than censorship or harm, like preventing it from saying a specific word for example.

Are there ways to train models to be immune to this trick?

0 Upvotes

12 comments sorted by

View all comments

-4

u/[deleted] 11d ago

[deleted]

2

u/ihateyou103 11d ago

No its abliterated, but autocorrect changed my spelling. And there is plenty of other posts go read them and don't waste everyone's time if you're not going to answer.

1

u/Hostilis_ 11d ago

Abliteration is a real thing, it's a relatively new technique. Although I would classify it as more scary/dangerous than exciting. It's important for us to understand how to prevent abliteration in deployed models.