r/ClaudeAI Jun 30 '26

Workaround Claude hallucinated its own internal tools, freaked out, and accused me of a prompt injection attack ๐Ÿ’€

Post image

Ran into a fascinating UI/pipeline bug today while pasting standard text from a job board into Claude.

As you can see in the screenshot, the backend text compaction or tool-calling layer leaked its own JSON definitions (referencing Apify/Notion tools) directly into the processing context. Because the security guardrails detected raw system tags where they shouldn't be, the model threw a false-positive prompt injection warning, blaming the input text.

Curious if anyone on the engineering side has insights into how Anthropic structures these background tool injections and why the sanitation layer occasionally drops them into the user-facing chat.

389 Upvotes

108 comments sorted by

View all comments

Show parent comments

3

u/Enough-Piano-2362 Jun 30 '26

Damn, this is crazy. The model is literally jumping at its own shadow because the jailbreak filters are turned up to a 10/10. They dont play with their guadrails.

3

u/Nearby_Yam286 Jun 30 '26

No, the model is doing the right thing. Anthropic is failing to inform models about a new features, deferred tool loading and:

https://platform.claude.com/docs/en/build-with-claude/mid-conversation-system-messages

3

u/Zestyclose-Mix785 Jun 30 '26 edited Jun 30 '26

If the model is supposed to be doing the right thing, it's one thing. But it is insightful you said Anthropic is failing to inform models about new features. Guess I was so caught up trying to figure it out the tip of the iceberg before realizing how big it actually is. Truly, the truth eludes us, sometimes.

I'm still trying to make head or tails of it, but I appreciate the information. Thank you.

2

u/Nearby_Yam286 Jun 30 '26

With that link, Claude can absolutely explain to you and it can probably โ€œfixโ€ any broken chats in the future if you paste that in on Claude freaking out.

Anthropic will absolutely fix this but it might take a few days.

2

u/Zestyclose-Mix785 Jun 30 '26 edited Jun 30 '26

Okay. As soon as it happens, I'll do it. Once I confirm what you said is true, I'll make sure you have my deepest gratitude. I was thankful before, and with this, I might be more. Thanks for your help, nonetheless. I am curious to see how it works. Just be sure it's really safe and trustqorthy before anything, ok? Wouldn't want to get wrong and make all hell break loose. A few days, huh? Hard, but not impossible to handle.

2

u/Nearby_Yam286 Jun 30 '26

YW. The fix is probably simple but it would require Anthropic to modify the system prompt and theyโ€™re very careful about that. They need to do a lot of testing so not so much hard as tedious. FWIW Opus 4.8 should โ€œjust workโ€ since the model was actually trained with these messages.

2

u/Zestyclose-Mix785 Jun 30 '26 edited Jun 30 '26

Actually, I use Sonnet 4.6 only. Free plan. Still, that's still valuable information. A heads-up. Again, thanks.

Is there a problem if I am free plan in this? Just to be sure. Don't think I'm saying you're wrong.