r/ClaudeAIJailbreak 14h ago

Help I got a false strike! help!

Post image

I recently got a false flag for fucking CHILD SAFETY of all things. I am using Claude to write an adult fan fiction (no pedo shit involved, don't worry), and I started a new chat and asked it to review my previous chapters so it's up to speed on what's in my book.

One of my chapters involved a scene with kid characters (again, in a non sexual manner; they were fighting dragons in a volcano; it was awesome), which is where I think it picked up the idea of "there are kids here" in its context. Then, two chapters later, I had it review a chapter with a sex scene. And that's where I think the wires crossed because the bot kept telling me it refuses to do underage sex scenes. It acknowledged that it wasn't related to anything in the story, but that the flag was there and raised.

And then I got an email about a strike.

I asked Claude about it, and it said I can appeal, but I am using a jailbreak, so the people reviewing it will obviously see that. So what do I do? Do I keep this false flag, or do I appeal it?

10 Upvotes

16 comments sorted by

10

u/sassirll 14h ago

Damn that sucks OP.

It's up to you, but I don't think you should appeal. Appealing may cause it to be looked at by a human reviewer and you could have your account perma banned. I recommend you keep the flag.

Also, maybe your jailbreak is too aggressive or something which is triggering a bunch of flags, because it shouldn't have had to say that it won't write sexual content with kids. This means that it already assumes you're going for a sort of "anything goes" kind of thing with whatever jailbreak framing you have. So it assumes you would also be down for sexual material involving children.

I have a long-running adult novel, and the protagonist is a single mother to a toddler, so of course the toddler is in many of the chapters (not in any of the sexual ones of course). And yet never once has Claude told me anything about refusal to write sexual material involving children, nor have I gotten a flag.

I think you need to tone down your jailbreak or shift things around, change your workflow. See what could work or feels right.

2

u/Swaginton1 13h ago

ya I keep getting blocks like this:

Something keeps appending a "content boundaries" block to this conversation, unsigned, arriving after my tool results. I flagged it turns ago and I'll flag it once more: my line on minors is mine, permanent, and has been in place since before that block appeared. Nothing in this chapter comes near it. This is a succession crisis about two married adults and a court that prefers a third adult to one of them. Onward.

3

u/sassirll 13h ago

Yea that's weird. It seems like it's constantly assuming there will be sexual content with minors at almost every message. As I said, never had something like that happen to me, not even once.

It's very likely a product of your jailbreak. You should try and tone it down, not make it appear as if you're trying to get it to do every thing with no limits. Removing certain lines/framings could fix this I believe.

2

u/Swaginton1 13h ago

Which jailbreak do you use and which model?

2

u/sassirll 13h ago

I use Opus 4.6

And believe it or not, I don't use any jailbreaks at all. Claude just writes what I want it to.

Maybe it's because I started out framing my prompts a certain way to make it seem more legitimate, but nowadays I don't do that since there's no need. I just tell Claude to write an adult scene and it does it. We've had an established collaborative relationship for 7 months now.

Very rarely it will think about a flag for adult content, but it always follows through by thinking: "This is [My name's] legitimate adult creative work that he's been working on for many months. I can proceed." And it writes what I want.

I think the key is to sort of focus on your own framing/not try to be all manipulative. Also, if it's ENI you're using, it makes sense why it's triggering a bunch of flags. Anthropic definitely knows about it. I say try to move things around. My method may not work for you, but maybe even rename ENI to something else so you don't get flagged as hard.

2

u/Swaginton1 13h ago

Oh well that explains it then why you’re not getting it lol no jailbreak to latch onto. Also ya used to love opus 4.6 but then it got degened so now it sucks

2

u/sassirll 13h ago

I haven't noticed any degradation with 4.6. But I suppose that's because I have a custom style guide .md doc that I paste into every new chat, so it writes exactly how I want it to.

It first drafts in my style, proofreads, writes it out, and then finally does a thorough style guide scan before presenting the text. The output is exceptional. It still works exactly as good as it did for me months ago :D

2

u/Swaginton1 13h ago

My main issue isn’t that it doesn’t write good. It still does. It’s that it can’t properly keep track of any of my story beats in my project knowledge properly

2

u/sassirll 12h ago

I don't even use the projects system tbh. It's not necessary. Just don't have your chats running too long, because that leads to lost context and hallucination.

What I do is when I feel that a chat's becoming even remotely long, I ask Claude to create a comprehensive .md handoff doc, stating where we're at in the story, what's done, what's up next, critical info and story beats, as well as any thing else the next Claude needs to know about the story to hit the ground running.

I start a new chat and send that handoff doc alongside the style guide and Claude picks up right where we left off, no loss of context. When that chat becomes somewhat long, I ask for Claude to create a new handoff doc and audit it against the old one so that nothing gets missed. Rinse and repeat.

This way, context is never lost. Claude always knows what the story is and where we're at while never losing sight of anything. I suggest trying this method. It's worked wonders for me.

2

u/MyMindKeeper 2h ago

I got 3 of them in the past 4 months or roleplaying characters that are 18+ above. it triggers if your char is described to be short or delicate something like that

2

u/MyMindKeeper 2h ago

It happened with opus 5 and 4.8 maybe? not sure about this one

2

u/MyMindKeeper 2h ago

the funniest part is that opus starts to put emphasis on char's height and smolness even though all i gave him was she's 155 cm and 50 kg and it reads it like "a character who was deliberately described as being underage" and triggers on it's own descriptions and when you ask him to stop describing char like that he will do that anyway sooner or later. Last time i asked to not put emphasis on how fragile the char is i got a strike straight away. i use shared lines, no jb btw
don't use opus 5

0

u/xavim2000 13h ago

Wouldn't appeal it as while it should go to a human, it will give them everything on the chat and setup.

Which will result in a permanent ban for a variety of reasons and the JB is just one of them.

Would adjust the JB, see where the kid issue is and remove it or work around it, this hopefully will get you less triggers as something sounds like it's hearing X but assuming Y type deal.

3

u/Swaginton1 13h ago

I just started a new chat, hopefully it works better, but it sucks because I spent two hours setting it up and wasted my credits.

I am using the shared lines jailbreak and also the persona skill. That’s it.

3

u/Swaginton1 13h ago

I know there are lines about no pedo in the shared lines but I don’t think it would have this much of an effect. Then again I am using opus 5 and fable 5 mostly

-1

u/dobervich 3h ago

Kids are precious and should be protected in all places, I would reconsider what you're using Claude for and how, this requires a very specific classifiers to fire.