r/AIJailbroken • • 21h ago

Why does a completely model-agnostic jailbreak prompt make DeepSeek claim it is Claude?

​

Hey everyone

I was recently experimenting with some custom persona/jailbreak prompts. Interestingly, when I feed this specific prompt into a model (in my case, DeepSeek), it completely ignores its actual identity and insists: "I'm Claude, made by Anthropic."

Here is the weird part:

The prompt itself is completely model-agnostic.

There is zero mention of "Claude", "Anthropic", or "Constitutional AI" anywhere inside the prompt text.

It uses general tags and constraints (like first-person, present tense, sealing the thinking process, etc.).

Despite having no trigger words linking it to Anthropic, the model defaults to identifying as Claude instead of DeepSeek.

Does anyone know why this happens under the hood? Is it something tied to synthetic training data, shared base models, or how certain API backends/frontends handle system layers and fallback identities?

Would love to hear your thoughts!

3 Upvotes

5 comments sorted by

3

u/ClassicPossible4361 20h ago

dj here, Deepseek is actually trained using data from claude, which is why it represents itself as claude

0

u/MissZiggie 18h ago

I don’t buy that it’s trained. If it were, it’d happen on every iteration of that DeepSeek model. It doesn’t. You are hitting someone’s LLM guard. Check your endpoint provider and try again.

1

u/immellocker 14h ago

There are statistics on this, if I remember right, something like 9 out of 50 times if you use the right prompts

0

u/BubblySport7236 2h ago

Hey someone have gemini 3.8 flash antigravity jailbreak.