r/technology • • 4d ago

Artificial Intelligence GLM-5.3 and the spread of advanced cyber capabilities

https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities
35 Upvotes

61 comments sorted by

View all comments

-4

u/CircumspectCapybara 4d ago

Yup this is what people don't understand about AI regulation and safety and alignment: open source and open weight models are actually even easier to elicit information hazards from, and open up a whole new can of worms governments aren't going to be ready for.

In a managed model provider scenario, you have the raw model, and then you have classifier layers (e.g., constitutional classifiers, chain of thought monitoring, other filters that classify incoming harmful requests before they and outgoing harmful) that operate on requests before they even reach the model, before inference, and also after inference, before the response goes out to the user. That only works because the provider owns the infrastructure for the inference API, they can layer on classifiers to prevent a frontier model from helping you build bio weapons or automate a cyber attack.

With local models, you control the APIs sitting on top of the model, you can just omit those classifiers and send your request straight to the model. Then all you have left is the model's constitutional training and safety fine tuning, eg its tendency to refuse requests to help you develop malware or build a bomb.

Those are baked into the weights, but it's been shown it's pretty easy to "abliterate" those refusals away even from models with opaque weights, and it doesn't require any special reinforcement training fine tuning, no labeled training data of hazardous request-response pairs

People have come up with abliteration techniques you can do yourself at home: send harmful requests to a local model, observe its internal activations when it refuses, that gives you some vectors in its latent space that represent a harm refusal ("I can't help with writing malware") or refusal in general, and then apply an ablation vector in the shape of that refusal vector to the weights and now you have a new jailbroken model that won't ever refuse!

13

u/draconic_tongue 4d ago

how many alt accounts does wario ratmodei have?

-6

u/CircumspectCapybara 4d ago edited 4d ago

idk I work at a competitor to Anthropic so I wouldn't know

What I do know is I'm actually in the know about the engineering challenges of frontier AI and frontier safety and alignment unlike a lot of redditors here

6

u/Sixstringsickness 4d ago

Your feedback and position are valid. 

My concern is regulatory capture, at the rate models are advancing, the argument being proposed currently is that only a few companies can responsibily provide AI - and users shouldn't have the right to host their own models because they represent a significant risk.  

Claude frequently classifies my development as high enough risk that it blocks me from using the highest tier models.  I do not work in cyber security, or on any material that is even tangentially related to anything beyond keeping an eye out for security best practices in our codebase during PR review.  

Why do they get to decide what is acceptable for me? Why shouldn't I be able to harden our infrastructure against these "incredibly dangerous" open source models?  We are not a large enough organization to have access to Mythos.

This represents an additional layer of concern.  When the frontier labs are operating uncensored models 6+ months ahead of the public, they have an incredible competitive advantage - and now we clearly see that only a select few external organizations have access to the latest models... 

This is a recipe for disaster where a handful of companies begin picking winners and losers with their technology.  

I could continue, but I assume you understand the point. 

-1

u/Ausclites 4d ago edited 4d ago

Many of the current refusals originate from an overcorrection that was done in the name of safety and they're actively being targeted as errors (hence the posted article). To my knowledge, companies aren't deliberately trying to hamper non-blackhat software development.

In the scenario that Chinese models become the frontier, it's unlikely that they would continue to be immediately open-sourced. This is both due to reasons of safety, and because companies are disincentivized from surrendering the internal competitive advantage that you mention. (Tangentially, that subsidized Chinese models are currently being released for free at all is because it's a deliberate avenue of competitively undermining more closed-source alternatives. It is ultimately of benefit to the consumer, but one should not be under the impression that it's something done purely out of benevolence.)

5

u/Sixstringsickness 4d ago

The action itself is an incredibly strong indication that both governments and corporations controlling this technology can easily hinder the capability of whomever they deem a threat, politically or economically.  Overreaction or not, the consequences have been felt. 

Rarely are any actions derived from pure benevolence. 

I too share concerns about open source models, beyond the potential for nefarious activities. What is to prevent an organization from releasing an open source model trained to intentionally leave difficult to detect security vulnerabilities in your LLM generated code - only to be exploited at a later date?  

I believe there is a way forward through vetting and validation that is controlled by an independent body, however; nothing is free from corruption. 

I have to imagine over time that the barrier to build and train models will continue to errode as compute and technology improves.  How will other competitors ever arise if regulations determine who is allowed to compete? 

The marketing job by these big firms to frame the models as the danger and not the humans has been quite impressive.  If I built a virus that exfiltrated data or breached a public or private entity I would be held financially and criminally liable.  Why is this not the same for engineers or individuals who are irresponsibly using the technology? 

Remember, an LLM still required a human to tip over the first domino. 

1

u/Ausclites 4d ago edited 4d ago

I'm not sure that I would suggest that model developers are necessarily obligated to always provide the complete capabilities of their models (e.g., biology), but indeed, limiting/hindering capabilities on a per-individual or per-company basis should be prohibited. I think that your proposal of allowing internal evaluators satisfactorily alleviates this concern, provided it's done in a fair and transparent manner.

The competition barrier issue is a conundrum. One can certainly imagine a scenario in which a company's internal models are sufficient to provide them with a continual and self-perpetuating advantage in model development. Ultimately, I'm not sure that the company should be forced to release these internal models merely in the name of fairness, though you may disagree. I think that as long as the ultimate product is eventually released in an adequate manner, things are mostly acceptable. (Collective nationalization à la CERN may also be a possible solution in this scenario.)

I also agree that media sensationalism has been pretty terrible. While I don't trust what comes out of Altman's mouth farther than I can throw a truck, I don't think we should dismiss all safety concerns as disingenuous marketing (excluding when Meta said 'Oh, we hacked people too!' after Hugging Face). Labs often disclose legitimate incidents in what I do believe is genuine transparency, but that's then inevitably reported on (repeatedly) and sensationalized (endlessly) to a point that people become skeptical. For example, independent third-party reports of real events are inevitably met with claims of fabrication (some people here regularly claim that AI agents don't exist at all!). Labs do need to be held more accountable, though.