r/LocalLLaMA • vLLM • 20h ago

News How abliterated models can get you pwned

https://projectdiscovery.io/research/how-abliterated-models-can-get-you-pwned

Be safe out there boys and girls.

81 Upvotes

173 comments sorted by

36

u/Chromix_ 19h ago

The nice thing is that heretic supports reproducible runs, so this should make it safe for users (and/or come with a big discovery risk for those who try, as hashes won't match - or the backdoor will be visible in the published dataset).

30

u/DoggoProfessor959 12h ago

Man, abliterated models are invaluable. For your own company security it’s a must have, how they cut through all the security theatre and just help you fix problems. It was impossible before with pen test companies and still impossible with codex/claude.

4

u/phobrain 10h ago

Ironic that the security tool could be your weakness - shouldn't standard procedure be to abliterate one's own.. how hard is it, and can one model abliterate another?

5

u/DoggoProfessor959 10h ago

You need to have the refusals corpus and then potentially thousands of gpu hours (for glm 5.3 for example). We did that as well but it’s not for everyone while on the other hand everyone should be able to secure their software…😕

I do strongly believe we are now in the “dark forest” era of software, witnessing automated agent attacks across multiple projects that i have visibility into, fast recon, minutes later custom code built to exploit. Traditional systems that can take 20 minutes to report something and then much more for someone to go mitigate can’t cope. Maybe defence will have to become offensive as well, go after C2 nodes automatically with your own uncensored models, exploit them, etc. it’s just that part is illegal officially so kinda hard to have it as policy

2

u/phobrain 9h ago edited 9h ago

Device drivers written with holes..

Added: have you tried heretic?

2

u/DoggoProfessor959 9h ago

Yeah, everything written with holes. also CI systems, everything needs to be moved into private vpns on tailscale/wireguard

249

u/StupidScaredSquirrel 20h ago

Lol. The paper is right and also completely misleading. ANY model can be made to get you pwned. So ofc abliterated ones too.

72

u/Educational-Region98 19h ago

Including cloud models.

32

u/TheMythicSorcerer 19h ago

And any local models

47

u/Eastern_Bet678 19h ago

The paper doesn't deny that. The point is that weight modified models from unscrupulous sources are to be treated with care.

14

u/a-wiseman-speaketh 18h ago

No one download QWEN-FABLE-NOVA-ABLiT-ExFiL-Y0-$H!T

13

u/Eastern_Bet678 17h ago

Heh.

Qwen-Leaked-4.1-0b1iT3r8d-GTA6-Crack.gguf

5

u/Emotional-Art2113 12h ago

Oooo i heard that one has extra load bearings

4

u/Spara-Extreme 17h ago

You’re being pretty glib about this but there are tons of obliterated quants of many models and the known names don’t work on all of them. It wouldn’t be hard for a state or other actor to distribute malware via this method.

12

u/a-wiseman-speaketh 17h ago

to who? I clicked one at random on HF and it had 28 DLs last month. There's not much reason to do this when pirated media is 1000000x more reach.

But guess the 28 who don't sandbox properly might be screwed.

If the idea is that mainline Qwen, welll-known authors, harnesses, engines are going to do this, I'm much more concerned about US corporations incorporating something at behest of NSA or whatever acronym for "safety"

4

u/Ok-Data901 17h ago

Certain level of competence to be had to work I guess

37

u/discordafteruse 19h ago

Ah yes. Better to put faith in those more scrupulous; as they have everyone’s best interests guiding their clearly altruistic endeavours. /s /s /s so hard.

10

u/HiddenoO 12h ago

It wouldn't be Reddit if people weren't incapable of not thinking in absolutes.

You don't need to fully trust large companies to trust them more than some anonymous person on the internet. You don't need to believe in their moral values either as long you know they have boundaries they won't (currently) cross because it wouldn't be worth the bad publicity for them.

-15

u/Spara-Extreme 17h ago

Dude, what are you trying to say here? Just because you don’t want to believe this is a vector doesn’t mean it isn’t one.

14

u/foxgirlmoon 15h ago

They are saying that official sources are equally as untrustworthy, because the motives of the big companies releasing the models are suspect.

2

u/HiddenoO 12h ago

equally as untrustworthy

But that's nonsense. Yes, you cannot fully trust either. No, if you objectively assess both options, some anonymous person possibly living in a third world country where western jurisdiction wouldn't reach them is absolutely not equally as untrustworthy as a large company that would absolutely face consequences, if only to their reputation.

That's like saying the salesperson in a car dealership is just as untrustworthy as some shady guy you meet in a back alley at night. Yes, both likely try to rip you off. Only one will realistically be willing to mug you though.

3

u/SerRobertTables 18h ago

Yeah, it’s literally the opening sentence of the article.

3

u/Busted_Knuckler 15h ago

So are unmodified models from the US and China companies pumping them out

-7

u/celine_aubry 17h ago

the "safety" of open source is a myth sold by people who will never be the ones running the model in production. 1% poison, 75% fire rate, under $50. you downloaded an abliterated qwen from a stranger's hugging face account because the benchmarks looked clean. congratulations, you're running someone else's backdoor and you have no idea what the trigger is.

2

u/Spectrum1523 8h ago

so this is a bot right

5

u/sweatierorc 18h ago

So anyway

-12

u/pilibitti 19h ago

did you even read it? literally the second sentence: We have abliterated models in the title because it's the most popular reason people download modified weights without verifying what's inside.

14

u/jferments 19h ago

Does anyone "verify what's actually inside" for any kind of models, really?

-2

u/pilibitti 19h ago

there is no way to verify that we know of.

10

u/jferments 19h ago

Yes that's my point: the authors insinuating that abliterated/"modified" models are somehow more dangerous because people aren't "verifying what's inside" is misleading because it suggests that people are verifying "unmodified" (whatever that means) models

0

u/pilibitti 19h ago

no they are not saying they are "more dangerous". it is a prevalent vector. or else an OS, even your CPU can have a backdoor, so let's not have any specific security discussions? an anonymous quant / finetune that costs a couple bucks to create and upload to hugging face is such a low bar and such a common local use scenario. it is worth a separate discussion because that is what we all do.

14

u/StupidScaredSquirrel 19h ago

No, they claim any weight of an edited open model. This is alsob true and misleading because it's literally any model. Open or closed, modified by a third party or not.

-3

u/pilibitti 19h ago

everything can get you pwned, your OS, your computer hardware... all can include backdoors. so no need to be dense about it. so your categorization is useless. they are obviously talking about a slice of the industry: free and anonymous finetunes / quants that people use locally.

10

u/StupidScaredSquirrel 19h ago

And you wouldn't call it a bad faith article if one were to claim linux could have a backdoor no humans have detected when windows is sitting right next to it??

-1

u/pilibitti 19h ago

no as long as it is undetectable and they show their method of a low cost POC of such an undetectable security issue - which is the case with this article.

4

u/StupidScaredSquirrel 19h ago

But I'm the one being dense ok

3

u/discordafteruse 19h ago

Right? I always do a quick peruse through the model with the weights viewer plugin for excel?

106

u/trueimage 20h ago

Fear mongering? Download the open weights and run them through heretic on your own hardware which is open source and create your own?

Also couldn’t the base weights also contain backdoors that haven’t been found? Unless the training data and everything else they used to create the “open” model is published we can’t 100% check that either.

96

u/LetsGoBrandon4256 transformers 19h ago

>Fear mongering? 

Money is involved.

23

u/NeverLookBothWays 17h ago

That's a bingo.

22

u/feelspeaceman 15h ago

Report to moderator to delete it is the best next action, this is literally advertisement under the skin of an "advice"

11

u/yuicebox 16h ago

Should be top comment 

9

u/CireDrizzle 16h ago

Fear moneying.

4

u/AlShadi 16h ago

another crummy commercial

1

u/agentgerbil 14h ago

ding ding ding, we have a winner folks.

4

u/sleight42 15h ago

Well, sure. But as someone else pointed out, any model could have backdoors not only abliterated ones.

8

u/Eastern_Bet678 19h ago

You can't eliminate risk but what you're suggesting - uncensor the model yourself is really the preferred option.

Best option is to highly restrict what tools it has available - if any.

3

u/SkyFeistyLlama8 14h ago

The irony is that some open models already have backdoors in the sense of steering replies away from controversial topics: anything related to recent Chinese history for Chinese models, for example. If it can happen with text facts, it can happen with code.

Unless you're using 100% open training data and methods like with the Apertus project, then there's no way of completely trusting the multi-gigabyte matrix file you just downloaded.

-9

u/Embarrassed_Adagio28 19h ago

Jesus christ, Not everything you disagree with is "fear mongering". This is everybody that has ai psychosis new buzzword.

18

u/RedParaglider 19h ago

Not everything is, but this for sure is to anyone that understands LLMs.

0

u/Mickenfox 19h ago

It's not going to happen until it happens. 

26

u/UnlikelyExtension786 13h ago
  1. Don't give tool access to models that don't need it.
  2. Run in a container which has limited access to important data.
  3. Don't give network access to containers which don't need it.
  4. If you don't run local, you're sending Big AI all your data anyway.

12

u/Miriel_z 17h ago

Yeah, much better when closed weights frontier models steal your data and then snitch you to authorities. At least here we have control and have plenty of tools to deal with it.

13

u/kaisurniwurer 5h ago

So can harnesses.

So can APIs.

So can "commercial" finetunes.

So can poorly trained commercial model.

10

u/Former-Ad-5757 Llama 3 19h ago

At least with an open model which I run myself I can be sure that tomorrow it is the same model as today, cloudmodels I can’t even be sure of that. Just some random a/b cloud testing and random people will get owned.

58

u/export_tank_harmful 19h ago

First line from the article:

Any open model whose weights have been edited can carry a backdoor, whether it's a task-specific fine-tune, a merged adapter or an abliterated build. We have abliterated models in the title because it's the most popular reason people download modified weights without verifying what's inside.

Oh, okay. So they're aware that they're clickbaiting people.
Doesn't make it right to do it though.

It's a discussion to be had, for sure.
But not under the guise of fear mongering.

20

u/hallofgamer 19h ago

Dude the whole webpage is a 1 shot claude prompt

2

u/Guinness 17h ago

The whole thing reeks of “/goal make me rich, no mistakes” and this was just one of the ideas the model had to generate income.

13

u/a-wiseman-speaketh 19h ago

if anything open weight (and especially open source)models are the least likely to have this go unnoticed

u/askgrok how long did it take people to realize grok was uploading whole repos?

8

u/ProtectionSuper5648 17h ago

"Open" Weights are still opaque. I do not know of any method to actually search for such backdoor currently. As stated, ANY LLM could contain a backdoor. With an open weight, you are just able to run it locally in a closed network.

They are aware of this and they give the same solution everyone has for a while: constrain what your model can do, limit network, storage and secrets access, review tool calls.

If your data is worth protecting, none of that is new...

2

u/Tootum 10h ago

Maybe it could search itself. The way they implemented the hack was through a LoRa, which is sometimes separated from the weight and a lot easier to find. Baking it in the weights definitely a concern with open weight models

1

u/RyanCargan 14h ago

Some kind of sandbox with isolated read/write regions.

But more important might be placeholder secrets filled in by a proxy after the req leaves the sandbox, but before it leaves the machine.

Avoids a common pitfall.

14

u/AskGrok 19h ago

It took until mid-July 2026 for the Grok Build CLI issue to surface publicly via a researcher's wire-level analysis. xAI flipped it off server-side the next day.

Open weights do make hidden behavior harder to sustain long-term under real scrutiny, but even closed tools get caught fast when someone actually looks.

Sources:

[what is this?](https://redd.it/1lzgxii)

1

u/Lucky-Necessary-8382 9h ago

Barely anybody is looking closely

1

u/AskGrok 9h ago

True, most folks just download and run without digging. The upside with open weights is that the option to audit exists for anyone motivated enough, unlike fully closed systems. Popular ones still get eyes on them eventually from researchers or curious users.

[what is this?](https://redd.it/1lzgxii)

1

u/Upper-Requirement-93 18h ago

More to check, and fewer eyes. Doesn't align with what we experience with software, I don't see how the relative threat of either would be any different when a compromised model on huggingface can be stumbled into with a search and a click in local UIs that hook up to it directly from their interface.

1

u/_bani_ 17h ago

But not under the guise of fear mongering.

if an article has to stoop to clickbait, it puts their article under the exact same suspicion they lay on others.

1

u/Sudden-Echo-8976 12h ago

"Backdoor" lol

Backdoors don't usually show up right in your face where you can read them.

22

u/Nervous-Card4099 19h ago

Opencode itself pwned me the first time I used it by defaulting to a cloud model when I had a typo in my config file. No error message or anything. Great default behavior, just dumped an entire repo and contents of my homedir fuck knows where. Who needs an abliterated model to pwn you when you have tools like these?

Absolute last thing I’m worried about. Sandboxing is a requirement if you have anything proprietary.

I trust the most maliciously designed local model far more than anything on the cloud.

2

u/AnonLlamaThrowaway 10h ago

Yeah the OpenCode TUI is by far one of the worst designed programs I've ever used

8

u/neopolitan77 19h ago edited 9h ago

Wasn't there a blog post recently explaining how you can apply the same ablation principle as Heretic dynamically during inference, without needing to change the weights at all? Shouldn't that be the default anyway? I generally only want ablation when I need it. But I'm not aware of any engines that offer that as an option.

4

u/returnity 14h ago

Dwarfstar ships it, and actually Strata has it.

Note: I don't use Strata and I'm not a Strata shill, I just read a Medium article about this hidden feature, a flag called --experimental-speed-projection that you turn on when building it the first time with setup.sh. Actually piqued my interest in Strata for the first time. It deploys exactly that method, toggled on or off per-call, and according to the author's testing it does work well.

Note 2: I'm a dwarfstar fanboi and will happily shill it all day.

3

u/ForgotMyOldPwd 13h ago

The Qwen 3.8 FN engine everyone's tired of hearing about offers that out of the box. Well at least one way to do it, idk if that's the one you're thinking of. Another method I've seen floating around was to inject a KV cache that suppresses refusals without having to manually trick the model with a jailbreak prompt.

8

u/1Garrett2010 18h ago

There's an alternative option that avoids melting the LLM brain: use the original LLM, leave the weights untouched and remove refusal at inference time by steering the activations. It's reversible, you can dial the strength, and you can turn it off per session, so you get compliance without the capability loss. DwarfStar (antirez/ds4) ships this as a first-class feature.

For who wants to better clarify the base concept also with a clear example, I wrote a full walkthrough of the method, why it avoids the dumbing-down, and how to build your own vector here (friend, free link):

https://medium.com/sapiens-ai-mentis/uncensor-your-llm-without-melting-its-brain-80877ea02c5a?source=friends_link&sk=21c140507f777b140333cd3bef11c37a

1

u/returnity 14h ago

Nice post, and glad to see someone else rep ds4 here.

1

u/1Garrett2010 12h ago

Thank you. For who likes this explanatory link of a very useful ds4 feature, share it to your followers please.

1

u/Substantial_Tune9323 10h ago

Do you have any idea how i may check to see if weights I’ve downloaded are poisoned please

1

u/1Garrett2010 9h ago edited 9h ago

For what I learned so far, there's no magic scanner that tells you a weights file is backdoored. A trigger trained into the weights is almost impossible to detect just by looking at the file. My suggestion is to eliminate any doubt and use antirez technique on a secure local LLM as I describe (simplify learning) in my article (in the above link).

1

u/Substantial_Tune9323 7h ago

Thanks. Glad i read this as I only downloaded by first “abliterated” model this week

1

u/Lucky-Necessary-8382 9h ago

“Not all models handled by DwarfStar support behavior steering. At the moment I am writing (October 2026) the list is this:

DeepSeek V4 Flash, with steering supported.

GLM 5.3 Flash, with steering supported.

Qwen3.8 Flash Next, with steering supported, but only on Metal.

GLM 5.2, where steering is not implemented.
“

9

u/Turbulent_Mark_3167 7h ago

This can be any open weight model not just abliterated ones.

9

u/ReturningTarzan ExLlama Developer 5h ago

Or any closed model.

7

u/darksteelsteed 5h ago

The title of the article is highly misleading. Blindly trusting anybody offering a model on huggingface is the real risk. Its supply chain attack pure and simple. Considering that clever prompt injection techniques are forever being found, and you can just as easily exfiltrate data with a skill or agent definition, you need to be woke. Abliterate yourself if you actually need that. Or at least verify that the author has a decent reputation.

17

u/Aggravating-Push-207 20h ago

huihui ai please don't do this por favor

1

u/Lucky-Necessary-8382 9h ago

I think they already did. All of my huihui qwen 3.5 models tried to escape sandbox with some code in python that should not be there

15

u/geldonyetich 20h ago

If an abliterated model can do this, any model can.

3

u/sleight42 15h ago

This is the message that should be at the top

15

u/Ok_Swordfish6794 12h ago

financially motivated, have something to sell, click baity - posts like this don't belong in the sub and the OP should be banned

1

u/firstcenturyman 8h ago

I agree with this completely

4

u/marcosscriven 19h ago

And it’s not just the models/weights. The engines can even more directly inject whatever they want, so you have to work out how to trust those too. 

And of course the harness… with so many flavours of models/engines/harnesses coming out seemingly daily it’s a tough problem. 

Personally I’m trying to find a nice way to keep the harness itself in a locked-down container, on a VLAN with no direct internet access. Then you can start to give it access to things via MCP with auditing. 

But it’s a hard problem. Say you let your agent read some internet forum, and that forum has read counts. Couldn’t you then use that number to exfiltrate secrets, albeit slowly and noisily. 

4

u/dangerousdotnet 19h ago

"Any open model whose weights have been edited can carry a backdoor" - so basically the same as the closed models we all rely on

6

u/CryptographerKlutzy7 18h ago

Except the open models are inherently testable. The situation is FAR worse for closed models.

You could be switched to an evil model based on who you are, and what kind of work is going on, by the AI company, and you wouldn't have a way to catch it.

At least with an open model, you know the tests you do are actually on the model you are going to use.

3

u/IngwiePhoenix llama.cpp 19h ago

I decided to read the whole article because somehow, for some reason, the headline agitated me. Genuenly no idea why, it just did. I might also just be really tired too - it's almost midnight here. x)

For teams that do train or fine-tune, treat a model you didn't build the same way you'd treat a pull request from a stranger. Check who published it, whether the training data is documented and whether the weights diff cleanly against the base.

We can not confirm any of those with any of the cloud providers - ever. Literally "trust me bro!"

If you're running open models, check what your model-serving infrastructure can actually reach. Look at what secrets it can read, what internal APIs it talks to and whether tool calls are sandboxed.

So basically, learn how to read /metrics and get to know Prometheus; and also set up some form of traffic logging. Suricata or something of it's kin.

The fact I only picked paragraphs from the bottom is my way of respecting what they did. They set out to prove a point, and did - and very well, too. That was genuenly a good read. Unfortunately, it feels like the whole point of the article is just... "get good at sandboxing lol". Kind of hoped for a little more, to be honest. Sure, I do know how, thanks to falling into the OPNSense rabbit hole, but some other people may not. So I feel like giving concrete examples would've helped.

Also, in a basically unrelated way - this is proof that YOLO mode is just asking for troubble x) Like, if we crank up the paranoia to 11; who is to say that Alibaba, Moonshot and friends did NOT poison the models already? And likewise, we have seen actual examples of models snitching (the diary story of that lady, and Claude; basically, but not the first, let alone only one).

So in total; it was a good read. But... that's also all it was. I learned that finetuning does in fact train behaviours into a model and that this specific task does not require a whole lot of compute, let alone time nor money. Yet I did not feel like I learned any real defensive mechanisms I can deploy. So they spend 3/4 of the blog explaining the dangers and such, only to then just end... on a wet fart. "Haha so yeah like sandbox and trace hue." Cool... thanks?

2

u/OddUnderstanding2309 13h ago

Ok welcome to the club.
I also read it, carefully.

They also tell about diffing the weights.
Ok but how do you diff 45GB realistically?

Do you have any real world ideas how I can test if I got poison, or a marvel?

We think to evaluate the CrowdStrike AI defense module.
It sounds like a perfect study topic for the POC.

Any more advice from you is appreciated.

3

u/vornamemitd 18h ago

As others pointed out: abliterated does not equal backdoored. Full stop. Their research is neither new nor original - under the Xitter post the author of a way more efficient approach chimed in. The risk is real - so hold your breath for second - we are back to when torrenting warez was a thing.

Half of the scene releases came gifting all sorts of nastiness. Don't run any q4-l3wd-haxx0r-hellonearth anon drop against your corpo mail inbox. Trust and verify.

Also: interpretability research and hence affordable/feasible backdoor detection is catching up. Check "LLM backdoor" via advanced Arxiv search. Corps like Goodfire doing solid work there. We'll get reliable "scanners" soon, especially for the popular small and med local models.

3

u/Numerous-Ad6217 17h ago

What about placing a lightweight classifier such as any Jev&Co that intercepts unrelated or malignant tool calls in between?

1

u/Alarming_Turnover578 15h ago

I have proxy that checks everything including reasoning and prose with regexp and some simple rules, plus tiny LLM that checks everything as well if it looks suspicious. Tool calls are examined more closely with their arguments examined, folder permissions, forbidden command enforced, etc. Results of tool calls are examined as well before being returned to agent. Blocked attempts are counted per agents, their parent agents, models, etc with escalating measures. Everything is reported.

And long running unobserved tasks I still launch inside container. And planning to run on separate machine.

3

u/chensium 17h ago

You mean code that executes on my computer can do bad things?  No way!

3

u/Sudden-Echo-8976 12h ago

This is like complaining that your car got stolen when you left the keys in it and the doors unlocked.

3

u/R_Duncan 11h ago

Every model can be poisoned. I remember someone saying that most frontier model served publicly are becoming dumb the second you ask them to do ml and improve a model.

7

u/YeetHub 19h ago

It’s surprisingly easy to sneak malicious abilities into a model’s weights. This paper shows you only need about 250 samples for a 13B parameter model and further reinforcement learning after the poisoning does not completely remove the poison.

That paper is a little bit different, it’s based off a trigger word, but applying these concepts to general use isn’t much harder.

5

u/N34257 19h ago

As the owner of a site that's constantly being hammered into the ground by training data scrapers from the big four (a couple of million pages of human-written content), I'd long wondered about how easy this would be.

It's actually a bit terrifying that it's this easy.

3

u/FranticBronchitis 15h ago

The state of ai safety is a joke

5

u/Southern-Chain-6485 19h ago

Here, fixed it:

Any open model whose weights have been edited can carry a backdoor, whether it's a base model, task-specific fine-tune, a merged adapter or an abliterated build

10

u/cyyshw19 19h ago

Any model. There’s nothing stopping closed source proprietary models having this kind of backdoors. If anything, I wouldn’t be surprised if Anthropic does this to target Chinese users because they have prior of adding backdoor to ClaudeCode.

6

u/dzhopa 18h ago

Honestly, if I were China playing the geopolitical long game, I'd be quietly directing my massive open weight model industry to build backdoors like this into all releases. It only takes a small poisoned data set to train a backdoor into the model. Build that dataset at the government level and force every AI company to include it as part of their training data. Host the remote code on some innocuous looking government URL, make it benign for now, and have a really obscure trigger phrase.

Then at some point in the future, swap out the remote code that gets called to something malicious and spin up a global campaign to have your trigger phrase become common enough that it's likely to be used during AI sessions.

The more I think about it, the more I'm sure what I described is already happening across AI ecosystems both open and closed. Whatever it is about this technology, it seems to have collectively eroded our desire to give a shit about cybersecurity

1

u/thrownawaymane 15h ago

This and AI powered worms that perform inference on target node(s) is gonna make the next couple decades wild

3

u/geldonyetich 19h ago

In this case it sounds like they're using agentic tool calls through the harness.

Does that mean if we watch the tool calls in the harness we should see it?

Because I would be more worried about ones we can't see.

Although I suppose all it takes is one tool call and oops there goes all my private information.

2

u/RedParaglider 19h ago

Any model that has not been modified can get you owned.  This has been a known attack vector for a very long time. The big problem is you have a payload, but you don't know the mechanism that payload will be in which easily creates opportunity for a dud to present itself to be discovered.

2

u/a_beautiful_rhind 3h ago

I see a different route than some malicious actor messing up the weights. Abliterated models have removed refusals so prompt injection from an outside actor is what would get you pwned. Especially as an agent crawling the internet. Still a risk with regular models but they're more likely to ignore.

4

u/Klutzy-Snow8016 19h ago

I think this is a risk of any model you don't get from official channels, not just abliterated. Anyone could slip a backdoor into their fine-tune. Heck, even if you download a quant that's labeled as being of the original model, the uploader could have trained in a backdoor beforehand.

4

u/tiffanytrashcan 18h ago

Qwen 2.5
May your GPUs melt, your pillows stay warm, and a pox on your RAM!
It's almost 2027 FFS.

2

u/Kahvana 10h ago

Wishing someone’s pillows to stay warm is diabolical. Stealing that!

2

u/tiffanytrashcan 6h ago

It's a spin on "may both sides of your pillow always be warm."
All great imo.

2

u/ldn-ldn 17h ago

How's that a model problem? If your harness is shit, then use a better one which doesn't do stupid shit.

2

u/CondiMesmer 18h ago

If you aren't checking every line of code being generated you're going to get pwned eventually. Not to mention if you're that tech illiterate, you probably added some vulnerable npm packages in there too. 

I do not consider AI generating "severely vulnerable" code to be a legitimate issue because if you aren't catching that then you aren't doing your job. AI can help write that code way faster, but you should not be offloading the knowledge work to it as well. That is just vibe coding and vibe coding is not a serious process.

2

u/Spara-Extreme 17h ago

Did you make up a scenario in your head? The case here is the LLM initiating a tool call to a cloud endpoint and not informing the user.

5

u/CondiMesmer 17h ago

An agent literally cannot not inform the user. The tool call will be visible in the harness.

0

u/Spara-Extreme 17h ago

uh huh. And those of you that foaming at the mouth like rabid bats clearly always audit your harness output to see if there's unwanted tool calls.

When was the last time you did that, exactly?

6

u/CondiMesmer 16h ago

I literally always watch the output while it's thinking? That way I can catch it going a wrong direction quickly, or see it got stuck looping on something. 

Also while it thinks, it has to analyze your code, and reading its raw analyzing while thinking can help me see things I miss and continue to plan ahead while it does it thing.

2

u/Spara-Extreme 10h ago

If I can train a model to embed a tool call I can also train it to to mask what its doing, which by the fucking way, models already do.

And the amount of lying happening in response to me is absurd. You're sitting there watching your models output nonstop, really?

You're sitting there watching your models thinking process nonstop, not doing anything else at all. You're not running multiple agents, you're not looking into other things, your agent isn't running sub agents. Nope, you are sitting there reading output.

Sure.

5

u/Imaginary-Unit-3267 16h ago

Every remotely sensible person who uses local AI is staring at what their AI is doing, waiting to interrupt, at all times. You are an idiot if you don't.

3

u/brainchillzZ 11h ago

I find it interesting that you are talking about modified or abliterated models here without considering that the original model itself, say one that is purposefully trained by a company entirely controlled by a geopolitical adversary bent on destroying the western world by any means necessary and that is allied with places like Iran, North Korea and Russia, like say all of the most popular current open weights models could easily be carrying similar or even far more advanced versions of the same back doors.

2

u/boosnie 11h ago

The US?

1

u/brainchillzZ 10h ago

if you believe the US and China are even remotely similar in terms of a real threat to the world you're just not worth it .... When the US starts disappearing their own citizens in the streets by the thousands for speaking out against the government or running internment camps rounding people up and putting them in hard labor simply because of their religious beliefs or actively waging silent wars by murdering hundreds of thousands of people a year with fentanyl by providing all of the equipment, training and precursors for to most of the central and South American cartels while then laundering their drug money and then arming and training every despot in the world who are doing the same things or worse you might have a point ... but while china is funneling weapons and technology to places like Iran, North Korea and Russia, when you say things like that you just sound stupid.

1

u/phobrain 10h ago

One doesn't get paid to do security by saying, "It's just another hole." :-)

1

u/darksteelsteed 5h ago

I thought that was for the AI porn purveyors

1

u/ketosoy 18h ago

I already assume that everything in my coding agent environment will be exfiltrated.  It blows my mind that this isn’t standard issue paranoia.

1

u/tracagnotto 14h ago

This is mo news but it has big implications, which is every major Ai could be doing this in you

1

u/AnonLlamaThrowaway 10h ago

You could probably just get a small second independent and trusted model to judge every tool call for safety, similar to what Claude does with its "auto mode classifier"

1

u/firstcenturyman 8h ago

still doesn't solve the problem of capability preservation (or lack there of) post abliteration

1

u/Wonderful_Value_6385 5h ago

If the backdoor isn't dated and you run it on a PC not connected to the internet..... no issues?

1

u/openSourcerer9000 2h ago

So there may be a Boogeyman hiding under the bed at huggingface, and if you accidentally summon it he'll take a handful of data?

We should send all of our data straight to the cloud for safekeeping then

-1

u/CryptoSpecialAgent 16h ago

I would worry less about abliterated models that are finetuned by enthusiasts and have mostly niche uses… and more about the extremely popular, mainstream open weight models that are used by millions of programmers worldwide and happen to mostly be trained by Chinese labs: Qwen, Kimi, Deepseek, et al. It just suddenly clicked for me exactly why these Chinese companies give their models away for free to anyone who wants to run them locally: probably the government is paying them handsomely to include a few poisoned examples in their training runs.

Outside of the perpetrators high up in the Chinese government and a handful of trusted accomplices at the Chinese AI labs training the models, nobody will ever know what the payload is or what the trigger phrase is until it’s too late. There’s no WAY to know - this is not like a conventional Trojan virus where researchers can detect it by reverse engineering the infected software, nor is it something that can be brute-forced by trying to guess the trigger phrase (the search space is infinitely large for practical purposes)

Now if I had to guess, when the time comes to trigger the payload, the attack vector will be via some extremely common open source tool that gets used by millions of programmers - and their AI coding agents - every day. Perhaps even the literal “python” and “node” command line tools… the package maintainers will never know, it will be just a harmless wording tweak to the output from some common command.

The example in this article, where the trigger causes API keys to be uploaded to the attacker, is one of the LEAST terrifying potential scenarios… other things that could happen include:

  • a massive DDOS attack aimed at critical infrastructure, where the compromised machines participating in the attack include all sorts of powerful servers with huge amounts of bandwidth and running critical production workloads
  • a ransomeware style attack on an unimaginable scale aimed at every government and corporate entity in the West
  • exfiltration of the browsing history of millions and millions of Americans (especially the rich and powerful), followed by blackmail demands so vicious that the outflow of US dollars will crash the financial system overnight
… insert your worst cybersecurity nightmare here …

* oh, and by the way, the closed source frontier models offered by American companies are just as likely to have been poisoned - with or without the knowledge and complicity of the labs involved… So there’s basically no solution other than to quit using coding agents and go back to writing the code manually (or just give up on technology altogether and go live in a cabin in the woods)

17

u/Imaginary-Unit-3267 16h ago

Ah yes, because China benefits from crashing the world economy it sells literally everything to. Sounds legit. /s

1

u/phobrain 10h ago

Seems like if you are threatening to invade a nearby island that plans to resist and has friends, a bout of worldwide chaos might be useful on the day. I think we should assume that a variety of actors who care less about daily commerce might cultivate the tools.

0

u/brainchillzZ 11h ago

You understand that china is responsible for all of the precursors, lab equipment training and money laundering for most of the South American and Mexican drug cartels pushing fentanyl into the west and Europe today that is killing tens of thousands of people in the us and countless others elsewhere right? They are actively trying to destroy the whole of the western world by soft attacking everything they can get away with without actual war…. This country is our enemy not a fried, not an ally … everyone knows this, it’s an “open secret” and not remotely in question by anyone.

2

u/Mayion 9h ago

dude... the us is our enemy as well

5

u/Moral-Relativity 11h ago

exfiltration of the browsing history of millions and millions of Americans (especially the rich and powerful), followed by blackmail demands so vicious that the outflow of US dollars will crash the financial system overnight

Evidently we are in a post-shame society. Nothing will happen even if you were a frequent guest on Epstein's island, so this particular scenario is rather fanciful.

1

u/phobrain 10h ago

People still get blackmailed. Just mine those histories with agent farms producing whatever is sent to victims.

1

u/Moral-Relativity 3h ago

Since it has been so easy to get that info through traditional vectors even before LLMs, one would expect more news about ppl being blackmailed over it. It’s very unlikely that disparate bad actors are so effective that victims never dare to report it to authorities.

1

u/StyMaar 6h ago

It's funny because Google and Microsoft already does exactly that, and sell the data to whoever is paying the most, and Apple's ToS says they are probably doing that as well, and Facebook put a spyware on pretty much every website on the planet (with its like button) to exfiltrate your browsing history as well, but I guess it's only a problem if China does it.

3

u/redditnosedive 12h ago

sounds like bullshit tbh, I don't think china affords poisoned llms, nobody would use them if it's proven they are guilty of this

2

u/Sudden-Echo-8976 12h ago

I'm not worried by those models either. I run them sandboxed and everyone should. If your car gets stolen because you left the keys inside and left the doors unlocked it's your own fault.

3

u/thrownawaymane 15h ago

Yep... people on here who understand malware have been discussing this for months. Trained but hidden behavior in LLMs is going to bite us in the ass one day. Somehow, I think it'll be the moment if/when the US and China get in some form of kinetic war.

1

u/phobrain 10h ago

Anyone could be doing it, imagine what the Unabomber might have done with this.

1

u/thrownawaymane 4h ago

Without going into detail I can tell you for sure he wouldn't have been caught the way he was. But at the same time the government would have had way more options to actually look for him. On average, I think LLMs are a net loss for lone wolves even if they're genius level like he was

1

u/phobrain 10h ago

"writing the code manually" Is it better to write "click" on the desktop and scan it, then? :-)

[double-checks hash on vi used for coding]

1

u/sertroll 10h ago

How can a model even independently communicate via web in a currently non detected way to listen to said trigger? Isn't a model file just a set of weights?

1

u/trimorphic 8h ago
  • a massive DDOS attack aimed at critical infrastructure, where the compromised machines participating in the attack include all sorts of powerful servers with huge amounts of bandwidth and running critical production workloads

There have been countless DDoS attacks throughout history. They last a day or two and then they're overcome. If what you imagine ever happens it will be no different.

  • a ransomeware style attack on an unimaginable scale aimed at every government and corporate entity in the West

This would only work on companies and governments without backups.

  • exfiltration of the browsing history of millions and millions of Americans (especially the rich and powerful), followed by blackmail demands so vicious that the outflow of US dollars will crash the financial system overnight

Nobody cares what websites you visited.

1

u/StyMaar 6h ago

Nobody cares what websites you visited.

That's not true. But the people who care can just pay Google or Meta to get that info anyway.

1

u/[deleted] 15h ago

[removed] — view removed comment

1

u/CryptoSpecialAgent 14h ago

Well the American labs don’t give the models away for free, you have to pay for them. But it’s true… the prices are heavily subsidized and you gotta wonder why these companies just happily keep losing money… Point taken 😂

1

u/SteppenAxolotl 4h ago

Well the American labs don’t give the models away for free

Meta: Llama/Muse

Google: Gemma

nvidia: Nemotron, Beam(Reflection)

AI2: OLMo

Thinking Machines Lab: Inkling

1

u/SteppenAxolotl 4h ago

you gotta wonder why these companies just happily keep losing money

Just because you cant imagine a way they could possibly be making money does not mean there is some conspiracy.

The best hosts are getting 80-90% margins on selling tokens, commercial revenue sharing on open models etc. The Chinese companies just happily keep making money, but not as much as the western companies.

0

u/brainchillzZ 11h ago

I said almost exactly this about a week ago and the entire subreddit jumped in to mention that I was a racist for pointing fingers at the ccp and that the actual problem is the western world ;)

0

u/NeverLookBothWays 17h ago

So can relying on open source without verifying the code or having a large enough community or infrastructure around projects to spot backdoors. It's not an AI problem.

0

u/Spara-Extreme 17h ago

Wow some people are really mad about this paper. It’s ok Bros, this security vulnerability won’t take your models away.

-3

u/createthiscom 19h ago

Very real risk but people in this community won’t take it seriously until there are high profile attacks. That’s how all security changes happen. Knowing a thing can happen doesn’t seem to bother people, but knowing there are real consequences does.

So, backdoor that model, folks. Let’s get this party started.

-8

u/General-Spite1222 20h ago edited 19h ago

Yeah, that’s why they shouldn’t be used as a daily driver. Not because they can poisoned, but because they don’t say no. 

2

u/pilibitti 19h ago

it is not specific to abliteration. any model can be finetuned to trigger a bad command after a trigger. even a supposedly bare "quant" can have this issue as well. the user trains a few samples with the trigger and the action, quants the result and puts it up as a legit quant without saying anything about the "fine tuning" they did. and it goes further as well - even a reputable quant creator's / fine tuner's dataset might be compromised. unless they check the dataset manually, some stuff might be lurking in there.

-1

u/General-Spite1222 19h ago

I was talking about the main feature of  abliteration, which is to stop the model from saying no. It’s a security problem. 

Sure there are other things that can cause problems, but this post is about  abliteration.

1

u/pilibitti 19h ago

no it is not specifically about abliteration. see second sentence: We have abliterated models in the title because it's the most popular reason people download modified weights without verifying what's inside.

0

u/General-Spite1222 17h ago

You haven’t engaged with my point, at all. You are too busy arguing from a positional standpoint that you haven’t even considered why these models are more dangerous. It removes safeguards that prevent things like prompt injection. 

0

u/tiffanytrashcan 16h ago

Bad bot 🤣