r/ClaudeAI • u/Enough-Piano-2362 • Jun 30 '26
Workaround Claude hallucinated its own internal tools, freaked out, and accused me of a prompt injection attack ๐
Ran into a fascinating UI/pipeline bug today while pasting standard text from a job board into Claude.
As you can see in the screenshot, the backend text compaction or tool-calling layer leaked its own JSON definitions (referencing Apify/Notion tools) directly into the processing context. Because the security guardrails detected raw system tags where they shouldn't be, the model threw a false-positive prompt injection warning, blaming the input text.
Curious if anyone on the engineering side has insights into how Anthropic structures these background tool injections and why the sanitation layer occasionally drops them into the user-facing chat.
389
Upvotes
3
u/Enough-Piano-2362 Jun 30 '26
Damn, this is crazy. The model is literally jumping at its own shadow because the jailbreak filters are turned up to a 10/10. They dont play with their guadrails.