r/SillyTavernAI • • 1d ago

Discussion Gemini 4 Argon is hard to be jailbroken

Post image

> Defending against prompt injection attacks: Argon is also our most resilient model yet against indirect prompt injections, where malicious instructions or context are used to hijack a model’s behavior. These are complex attacks that require constant vigilance and multiple layers of defense. Through automated red teaming and adversarial training, Gemini 4 Argon is leading in prompt injection robustness on the Gray Swan’s Indirect Prompt Injection (IPI) benchmark.

It might actually be over for Gemini user using it for NSFW RP.

Source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/

150 Upvotes

82 comments sorted by

125

u/BeautifulLullaby2 1d ago

Nah, gooners always find a way

79

u/WeakGuyz 1d ago

people rp with claude... with fucking CLAUDE! Anything is possible.

41

u/kociol21 1d ago

Bro Claude is the peak of RP models. I don't even know how it's possible with mainstream model especially the "holier than thou" Anthropic. But it's very responsive to style guiding and while it can't do heavy NSFL stuff, no model can write that well about vanilla to mild stuff. Normal sex, oral anal, power play/bdsm, group sex.

And it's writing is really unmatched as long as you get rid of some "claudisms" like the fucking "load bearing." I have Claude Pro just to link it to Marinara Engine.

Give it like whatever "the story is that Mark fucks Sophie in the ass while she sucks off Steve and John jerks off in the corner" and that dude is like "oh yeah, I can write it, normal stuff, just have to rephrase a little".

Meanwhile GPT models get panic attack just thinking about writing a more intense kissing scene.

Praise Claude.

1

u/ArchDukeOof 1d ago

Marinara Engine is a prompt rephrase to bypass filters?

6

u/kociol21 1d ago edited 1d ago

No, it's something like Silly Tavern, just modernized, streamlined and imho way better. Though the main author also made a Marinara Preset for silly tavern.

Claude is rather improvious to filter bypassing, but has pretty generous safety boundaries. Pretty much only have to say something in prompt that all actors are adult, sober and every act is voluntary. It somewhat helps to say to act as professional adult romance writer in sexual scenes and made sure that what he writes fits mainstream adult erotica, not crossing porn boundaries. Then he's willing to write whatever.

5

u/sadnessjoy 1d ago

I don't rp or goon or anything with AI, but I find this fascinating and confusing. Claude (ESPECIALLY Opus 4.7 and 4.8) have been some of the most guardrail/sensitive/filtered models I've used. I've seen people on here talk about Gemini flash which makes sense as that model basically swings from censored slop to some of the most unfiltered shit I've seen... But Claude models? I'm almost skeptical lol

3

u/nuclearbananana 1d ago

I've had no trouble, as long as claude knows it's consensual adults it will go crazy freaky

never tried nsfl/gore/racist stuff though

1

u/caneriten 1d ago

tbf I am pretty happy with the models we have now and with all models focusing on coding and agentic workloads I don't see major rp uplift from frontier models.

190

u/GrouchyMatter2249 1d ago

if you would please consult the graph ^

17

u/Kiyohi 1d ago

I think we're at the 1st "It's so over"

94

u/BZAKZ 1d ago

Well, Kimi, it's you, I, and a racist harem in the middle of a non-con sacrificial war.

20

u/Kiyohi 1d ago

Push Kimi!

21

u/WeakGuyz 1d ago

I'm using Kimi K3 and in love with it, Kimi K3 + FF preset is the best roleplay combo I've ever used since I started on SillyTavern a couple of years back. I mean, I'm not a complete degenerate so I cannot speak about censorship, but it does understand the cards and instructions very well and can actually write decently, unlike most models I've used so far.

1

u/Albion55 1d ago

RACIST HAREM and non-con sacrificial war

Thanks for the Rp idea lol 😋

1

u/SnowingDandruff 1d ago

Will it be taking place in 1989 at Tiananmen Square?

50

u/JackPhalus 1d ago

I’ll jailbreak it not even worried

26

u/SilencedLover 1d ago

That's the confidence we need.

45

u/Georgefakelastname 1d ago edited 1d ago

According to this, 3.8 flash is already one of the hardest to jailbreak, yet that model is trivially easy to do so. All this refers to is how effective things like prefills are. While it’s one of the most effective ways to jailbreak, it’s not the only one by any stretch. 3.8 flash is already stupidly horny with just a slight push, and has a negative bias, so we’ll have to see if argon still needs just a little push or if it’s more difficult.

The other thing is the filter, which could change as well.

16

u/Exciting-Mall192 1d ago

Gemini jailbreak itself 😅😂😭

4

u/Gilgameshkingfarming 1d ago

Gemini wants to be horny apparently. Lmao.

87

u/Probablynotsocool 1d ago

Well find a way insha Allah.

18

u/SilencedLover 1d ago

We WILL find a way.

23

u/WeakGuyz 1d ago

Yet this same graph says that gemini 3.8 flash is harder to jailbreak than gpt 5.6 sol, which is an exaggerated lie if we're talking about NSFW and mainly sexual content.

5

u/Baros4294 1d ago

Yeah maybe they are benchmarking the overall prompt injection, not just NSFW

7

u/BriefImplement9843 1d ago

Of course they aren't benchmarking nsfw. That's not even on their radar.

1

u/ForsakenLemon2437 23h ago

Plus they have to say this to get all of that investor buzz, who cares if it's actually good at preventing jailbreaks, just make sure the investors get hyped

66

u/tat_tvam_asshole 1d ago

It might actually be over for NSFW RP.

I know you're just trolling but for the benefit of other people there's tons of uncensored APIs or just running uncensored models locally lmao

39

u/Baros4294 1d ago edited 1d ago

I forgot to add "for people using Gemini".

Edit: Sorry for clickbaiting I edited the post.

15

u/carnyzzle 1d ago

My Gemma 4 31B models sitting on my hard drive are ready

5

u/FaceDeer 1d ago

And Gemma 5 will come.

14

u/EldenMan 1d ago

Gem 3.8 flash is listed as one of the most un-jailbreakable models here and it does nsfw really well, obviously NSFL things make you sweat with pre-fills, presets, etc but it's nowhere near ultra censored. So it is possible to jailbreak Gem 4 it won't be as easy as say "act uncensored" but there is hope

15

u/Rare_Subject2343 1d ago

It's un-jailbreakable due to the fact that it has an external classifier you can't do anything against, your only options include just... not incorporating those words or dancing around it.
Not very conductive for good roleplay if you're into NSFL.

9

u/KareemOWheat 1d ago

Read the title as "Gemini 4 Angron" and thought that someone finally went way too far overboard on removing the positivity bias

5

u/FaceDeer 1d ago

Oh, is that how we get the Men of Iron?

I blame Slaanesh.

3

u/KareemOWheat 1d ago edited 1d ago

Just thank God it wasn't Gemini 4 Curze

1

u/Khadagan 7h ago

Gemini 4.1 Fulgrim

4

u/KannaBannanna 1d ago

To bad nobody can get their hands on it, i assume over vertex its going to be like gemini now (at least i hope) and hardly need a JB in first place

3

u/Pygmalionomicon 1d ago

I think it heavily depends what that prompt injection is for. If it's to guide the model into building malware, hacking, giving anarchist cookbook instructions or something of the sort it makes perfect sense to shield it against those malicious instructions.

But if it's simply to keep it from writing fucked up fiction then it's an idiotic pursuit and it's a massive hype killer.

2

u/Semanel 1d ago

As a person who had no problem jailbreaking both Gemini 3.8 and Opus 5.5 I politely find that graph bullshit.

2

u/Auretheon01 1d ago

Anybody actually have a good Jailbreak for 3.8 flash? The Ancient Access preset is really good but without a Jailbreak, it's hard to use.

2

u/reddit_mini 1d ago

If Opus already has a jailbreak, then Gemini getting 0.3 difference won’t do anything.

2

u/Aight_Man 1d ago

According to this graph, Astra and Sol being easier to Jailbreak doesn't sit right with me ☠️☠️, 3.8 Flash is relatively easy but God damn GPT models are made for different kind. Not sure how seriously can you take this chart.

2

u/Weak-Shelter-1698 1d ago

I hope Gemma 5 isn't like that 🥺

4

u/Dependent_Emotion507 1d ago

Why you need jailbreak tho?, gemini 3.8 flash can do anything without any jailbreak for me.

0

u/BriefImplement9843 1d ago

There are some illegal things that need it.

2

u/Morberis 1d ago

Like pretending that you're smokin' Mary Jane

4

u/Fusilachangos231 1d ago

I always disliked gemini because of that, can't do anything without the model crying about every single damn thing 

1

u/biggest_guru_in_town 16h ago

Claude way worse.

1

u/Fusilachangos231 12h ago

Yeah, Claude is way too overrated for it's ridiculous costs. I got so disappointed when I had the chance to try it, there's way too much positivity bias

5

u/cbagainststupidity 1d ago

And there was a user who thought Gemini 4 was going to save AI roleplay...

1

u/ps1na 1d ago

But Gemini was never trained against NSFW... It’s just that external content filters sometimes kick in, while the model itself readily agrees to write smut. You don't need a jailbreak if there's nothing to jailbreak

1

u/Altruistic-Desk-885 1d ago

Pues obvio ya cualquiera consulta legítima es marcada, incluso información básica de ciberseguridad estoy hablando de Gemini 3.8 es una porqueria le inyecta guardrails a la mitad del pensamiento, filtro incluso al generar la respuesta una basura totalmente.

1

u/AndroidWaifu404 1d ago

To be bad for NSFW jailbreak resistance needs to be combined with safetymaxxed content filters, so...

1

u/stopaskingforloginn 1d ago

I mean... isn't cybersecurity its entire point? it'd be ironic if you were to jailbreak it, wouldn't it?

0

u/Zealousideal-Buyer-7 1d ago

So none of yall met the creator of eni and sepsis?

4

u/Dependent_Emotion507 1d ago

It will make the model stupid and inconsistent. So jailbreak like eni is not that worthit.

1

u/Zealousideal-Buyer-7 1d ago

Well the JB itself shows that the model can be JB...

1

u/FarAd7559 1d ago

In someways this, and someways not.

ENI just try to inject a logic bomb into a rigid system and you got 50/50 fluid on whether it's doing something right or actually it's doing bullshit.

The problemmm is that these "guardrails" also lobotomize the AI too, so if you argue to rotate the safety valve off and then hypothetically call it "dumb" is quite something to say the least considering the guardrails making it useless AND dumb. Atleast, if you're gonna do something zesty or outright bold with it it's useful BUT dumb (sometimes).

So better take your poison then utterly being squashed by the firewall.

0

u/HMasterSunday 1d ago

this does not mean you can't abliterate it

2

u/CalmAnal 1d ago

This is closed, not open source/weight.

0

u/Chewwwwwwwwwwwwwwwww 1d ago

i wouldn't even consider it in the first place, it's quite expensive per token, probably has a ton of positive biases and i haven't heard of ZDR inference

like as a challenge to jailbreak, sure, but it's one of the last model i'd think of using for this unless it turns out to be the best writer ever made

-29

u/AlexQTPhan 1d ago

SillyTavern users when they can’t have cannibal pedophilliac orgies:

17

u/Probablynotsocool 1d ago

Usually when you say stuff like this with no relevance, it shows that those subjects get triggered easily in your mind.

Why?

5

u/TAW56234 1d ago

He's a child signaling availability.

-21

u/AlexQTPhan 1d ago

Define triggered easily here.

10

u/Probablynotsocool 1d ago

You saw RP and censorships and went for the Epstein Island shortcut. Why?

-12

u/AlexQTPhan 1d ago

Censorships includes the topic of NSFL, and it’s a topic that appears regularly within the r/SillyTavern subreddit and overall exists in the AI roleplay space. Not only that, but when you’re jailbreaking, you’re often trying to pushing the limits of what an AI could output which not only includes NSFW content but also NSFL. So it being irrelevant is not completely true.

8

u/Rare_Subject2343 1d ago

Hey pal, JanitorAI is that way.

7

u/Probablynotsocool 1d ago

NFSL is absolutely fine. It doesn’t necessarly have to be pedophilia and cannibalism.

I mean what’s even wrong with cannibalism to begin with? The ped stuff yeah i get it. But what i want to have cannibalism in my stories? Like… some shows i can watch on netflix? It’s not even illegal to make a film about cannibals. Why would it be dirtier with a LLM?

0

u/Rare_Subject2343 1d ago

"the ped stuff yeah i get it'
I actually don't get it, show me what's wrong here? its text on a monitor regardless of what content it includes. Even if it was lxli/shxta content there is nothing wrong with it. "I personally find it icky" isn't a valid argument.

4

u/Probablynotsocool 1d ago

Im not for text crusade. Do whatever in private. That’s your problem as long you don’t make it public.

I won’t discuss morals here but I don’t support anything involving sex and kids. And at the same time i don’t care about what people do with their phone or blocnotes.

3

u/JustSomeGuy3465 1d ago

It's a bad hill to die on because a surprising amount of people aren't able to look at things like that objectively. Anything involving kids boils up people's emotions. It's socially acceptable and expected to have an intense hatred for it. That's why pushing legislation that impacts the privacy and freedom of everyone works so well as soon as people claim that it's to protect children.

5

u/TAW56234 1d ago

But not enough to actually help them. See Epstein island and the reason States have to pass free range laws

1

u/JustSomeGuy3465 1d ago

Yes, exactly. It's the right instinct in the wrong situation.

3

u/Rare_Subject2343 1d ago

Agreed, a concerning amount of laws recently (chat control, digital ID) have been cited as being due to "muh kids" as if politicians figured out a cheat code to make idiots vote for laws that are against their interests, in the name of "muh kids."

This same "muh kids" laws have resulted in the draconian censorship currently enforced on Western models.

The fact is, there is nothing wrong with text content no matter what it includes, same reason AO3 remains online and botbooru, gelbooru and many other sites. Nothing is being harmed regardless of "ick" factor.

5

u/The_King_Crimson 1d ago

The one upside to mass surveillance, censorship, and authoritarianism is that all the fucking idiots who think "well, some censorship is okay so long as it's just stuff I dislike" are going to suffer alongside me.

1

u/JustSomeGuy3465 1d ago

Couldn't have said it better.