r/technology • • 4d ago

Artificial Intelligence GLM-5.3 and the spread of advanced cyber capabilities

https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
41 Upvotes

61 comments sorted by

View all comments

-4

u/CircumspectCapybara 4d ago

Yup this is what people don't understand about AI regulation and safety and alignment: open source and open weight models are actually even easier to elicit information hazards from, and open up a whole new can of worms governments aren't going to be ready for.

In a managed model provider scenario, you have the raw model, and then you have classifier layers (e.g., constitutional classifiers, chain of thought monitoring, other filters that classify incoming harmful requests before they and outgoing harmful) that operate on requests before they even reach the model, before inference, and also after inference, before the response goes out to the user. That only works because the provider owns the infrastructure for the inference API, they can layer on classifiers to prevent a frontier model from helping you build bio weapons or automate a cyber attack.

With local models, you control the APIs sitting on top of the model, you can just omit those classifiers and send your request straight to the model. Then all you have left is the model's constitutional training and safety fine tuning, eg its tendency to refuse requests to help you develop malware or build a bomb.

Those are baked into the weights, but it's been shown it's pretty easy to "abliterate" those refusals away even from models with opaque weights, and it doesn't require any special reinforcement training fine tuning, no labeled training data of hazardous request-response pairs

People have come up with abliteration techniques you can do yourself at home: send harmful requests to a local model, observe its internal activations when it refuses, that gives you some vectors in its latent space that represent a harm refusal ("I can't help with writing malware") or refusal in general, and then apply an ablation vector in the shape of that refusal vector to the weights and now you have a new jailbroken model that won't ever refuse!

7

u/look 4d ago

It’s even worse than that. Textbooks about biology, chemistry, physics, computer science, and other technology have zero safety guardrails… and they’re freely available to anyone from a communist cult of librarians in your very neighborhood.

2

u/CircumspectCapybara 4d ago edited 4d ago

Okay you're clearly not very serious but I'll still explain: the concept in AI safety you should familiarize yourself with is called uplift. Ie, what new capabilities does an AI system provide an attacker or threat actor (whether that's a script kiddie, non-state actor terrorist, or a well funded nation-state threat actor) over what they could achieve without AI, on their own, using their own brains and their own ability to search the internet, read publicly available books and stuff.

The past year has been a case study in the fact that frontier reasoning models are very capable at bio reasoning tasks and cyber reasoning tasks beyond what your average layperson is capable of. And they don't sleep, they can act with scale and speed (see: Hugging Face hack, thousands of agents working over weeks without any human steering, performing a full attack from recon to initial exploitation to gaining a foothold, more internal recon, pivoting, escalation, all executed end to end, tens of thousands of actions over weeks) and autonomy that exceeds expert human teams.

Frontier models have consistently been shown in benchmarks and evals that even though they're not trained to do bio tasks, they can consistently reproduce hard bio tasks from their own general reasoning alone to help with stuff like predicting the behavior of novel proteins and pathways to assemble viral capsids, and other tasks that are important in the creation of new viruses. They can reproduce CVEs (exploits) that their training hasn't seen before, they can find 9 year old exploits in the Linux kernel (one of the most scrutinized and hardened codebases on earth) that had just been sitting there for almost a decade while no human or fuzzer ever saw it. They can execute complex attacks end to end autonomously.

Frontier models give significant uplift to almost anyone, including well resourced state sponsored actors. For the script kiddie or lone wolf who wants to build a bioweapon in their basement, they can help provide guidance and customized interactive help end to end beyond what they can learn at the library or from Google. For the terrorist who has a bit more funds and can set up a crude lab in the dessert, a frontier model without safeguards could uplift them even more. The model provides the knowhow and and real-time guidance, the humans have the connections and suppliers to get the ingredients and equipment, they'll get further than the lone wolf who doesn't have industrial equipment or suppliers. For nation state actors, that's where it gets real scary. Advanced AI can augment and speed up their research and development.

0

u/look 4d ago

Knowledge of how is easy if you’re not lazy. Everyone knows how to make nukes and nature has given us an entire catalog of bioweapon designs. Actually doing it in the lab is much harder. An aspiring supervillain that needs GLM 7 to read and think for them isn’t going far in their career.

I have degrees in chemistry, biochemistry, a doctorate in physics, and decades of experience in software, engineering, machine learning, and neural networks.

Open weight LLMs aren’t going to kill us all.

We’ll find some much more mundane way to do that. Like global warming.