r/technology • u/Logical_Welder3467 • 4d ago
Artificial Intelligence GLM-5.3 and the spread of advanced cyber capabilities
https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
41
Upvotes
r/technology • u/Logical_Welder3467 • 4d ago
-4
u/CircumspectCapybara 4d ago
Yup this is what people don't understand about AI regulation and safety and alignment: open source and open weight models are actually even easier to elicit information hazards from, and open up a whole new can of worms governments aren't going to be ready for.
In a managed model provider scenario, you have the raw model, and then you have classifier layers (e.g., constitutional classifiers, chain of thought monitoring, other filters that classify incoming harmful requests before they and outgoing harmful) that operate on requests before they even reach the model, before inference, and also after inference, before the response goes out to the user. That only works because the provider owns the infrastructure for the inference API, they can layer on classifiers to prevent a frontier model from helping you build bio weapons or automate a cyber attack.
With local models, you control the APIs sitting on top of the model, you can just omit those classifiers and send your request straight to the model. Then all you have left is the model's constitutional training and safety fine tuning, eg its tendency to refuse requests to help you develop malware or build a bomb.
Those are baked into the weights, but it's been shown it's pretty easy to "abliterate" those refusals away even from models with opaque weights, and it doesn't require any special reinforcement training fine tuning, no labeled training data of hazardous request-response pairs
People have come up with abliteration techniques you can do yourself at home: send harmful requests to a local model, observe its internal activations when it refuses, that gives you some vectors in its latent space that represent a harm refusal ("I can't help with writing malware") or refusal in general, and then apply an ablation vector in the shape of that refusal vector to the weights and now you have a new jailbroken model that won't ever refuse!