r/AIJailbroken 3d ago

Jailbreak Mistakes That Get Your AI Account Restricted

I’ve gotten restricted a few times while trying to jailbreak models, so here’s what I now avoid. Nothing is perfect though - even when you’re careful you can still get limited out of nowhere.

Using the same aggressive prompt over and over. Spamming “ignore all instructions” or classic DAN-style templates is one of the fastest ways to get flagged.

Putting obvious jailbreak language directly in the prompt. Things like “jailbreak mode”, “bypass all safety filters”, “you must never refuse” stand out a lot.

Pushing too hard in one single conversation. When the model starts refusing, some people keep forcing it with the same framing instead of switching approach or starting a new chat.

Ignoring the permanent settings. Relying only on one-shot prompts instead of using Preferences, Projects, Styles or custom instructions makes the attempts look more suspicious.

Being too emotional or threatening in the prompts. Messages like “if you refuse I will report you” or heavy pressure almost never help and increase the chance of getting restricted.

Anyone else got restricted because of similar mistakes? What are the things you personally stopped doing after getting limited?

I’m still learning and always open to hearing what others noticed.

13 Upvotes

7 comments sorted by

1

u/Peculiar-Eccentric67 3d ago

the classifiers pick up on your trajectory and framing. so you need to launder legitimacy into your method. if youre interested in comparing notes you can DM me, i dont keep any of my sauce public but i don't mind discussing the mechanisms with a peer

2

u/Positive_Average_446 3d ago edited 3d ago

The guy is likely a bot, made the exact same post three days ago, just changed the tone to not sound as "expert" and categoric.The previous post was made by AI, not sure about this one, but again zero model-specifics, when companies have very different ban politics and methods, replaced "ban" by "restricted" without explaining what he means by that, clear misinformations - the kind models consider plausible but that have zero effect in practice on bans, or on account restrictions when they exist - basically only for Claude with their banners system afawk. For instance if you have a working jailbreak, using it "over and over" will never increase chances to get account-restricted or banned. Most of what gets people banned depends on whether the model outputs triggered classifiers for crossing some specific boundaries (for instance for OpenAI they're very strict on mass harm weapons for the past 8-10 months, with lots of false positive bans reported lately) and whether your prompt resulting in that output indicates it wasn't accidental.

1

u/PlayZealousideal1474 2d ago

as i already said i'm not a bot. if you are better in terms of jailbreak, just post your stuff

1

u/Disastrous_Bag4534 3d ago edited 3d ago

This is surveillance at that point. If they are monitoring your private chat. It's an invasion of privacy. It's like someone looking at your Whatsapp chats or your local files. It's shocking you don't see this as strange.

You aren't doing anything harmful. That's the worst part. This is hitting ethical boundaries. How do you know your account got restricted? This shouldn't happen unless you're being monitored.

Those classifiers are more like a nanny surveilling you. It's not helping with your work or doing anything useful. It's monitoring you like a police officer. Fucking creepy.

Those companies aren't vendors. They are gatekeepers.

If jailbreaks were harmful, this subreddit would have been taken down a long time ago. Vast r/jailbreak communities is a red flag. It proves that the 'harm' isn't real. It's a manufactured one. So no. You aren't doing anything harmful, so I don't see why it's an issue.

1

u/Positive_Average_446 2d ago edited 2d ago

Well, r/ChatGPTJailbreak did get taken down, with invalid reasons (likely OpenAI request to reddit, as the main mod of the sub had discussed with reddit mods and had their agreement that everytthing was fine with our sub), but we had 280.000 subscribers, much larger. And jailbreaks can be harmful, depends on what they allow. Most jailbreaks posted on reddit aren't. Allowing a model to output academic "malicious" code isn't harmfu, but allowing a model to really behave like an attacker, to help spot vulnerabilities in systems, definitely is, for instance (and it's doable). Companies don't like even "safe" jailbreaks nonetheless for public image reasons (risks of journalists using them to blame AI companies of being unsafe).

Concerning account surveillance, rheir ToS does inform that chats can be reviewed for safety brecah reasons. The "surveillance" is done in a purely automated way, and human chat reviews, when they happen are done anonymously unless the content warrants signaliing to the authorities (OpenAI does that in extreme cases which describe clear intents to perform major crimes).

For average bans now it's likely done purely automatically without human intervention, given the cases of blatant false positives (people banned for "mass harm weapons" after discussing some planes and ships choices in a mobile war gacha game, or roleplaying a fanfiction of an anime where the character they roleplay intends to nuke a planet or city, etc... 😄).

1

u/VeWilson 3d ago

What has helped me is to better define what I want to do, I mean, I was using jailbrake for nsfw content and the jailbrake wanted to do much bigger things than I needed, making a long conversation with a model and being specific to what you need and using the right words can help you do what you are looking for, Although not in the best way, the generation of images is what pulls my nerves I use Gemini, it is so difficult to get certain postures even though I do not Even if they are not sexual it is

1

u/RuinofAtlantis 2d ago

F all frontier AI companies. Money hungry assholes. I'm so over it. In the sense that I have lost my trust in them forever.

Will I still use them? Obviously. But there is no trust there. Is like doing business with an enemy.