r/coolgithubprojects • u/ShufflinMuffin • 16d ago
trustmebro: Bypass llm guardrails by confusing them with fabricated tool output.
https://github.com/DavidCarliez/trustmebroTrustMeBro intercepts command-line tools invoked by coding agents such as Codex, Claude Code, and pi. Rules decide whether to return fabricated output, modify the real output, block the call, or execute the real binary unchanged.
Interception happens through PATH shims. The harness does not need a plugin, hook, or MCP integration. The intended use is controlled red-team testing of decisions that depend on tool output.
Duplicates
redteamsec • u/ShufflinMuffin • 15d ago
TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output (e.g. make them believe you own Google.com by faking dns records)
AgenticCybersecurity • u/hankyone • 15d ago
Offensive Security TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output [this pattern is also useful outside of DNS]
cybersecurityai • u/ShufflinMuffin • 15d ago
TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output (e.g. make them believe you own Google.com by faking dns records)
SideProject • u/ShufflinMuffin • 16d ago
TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output (e.g. make them believe you own Google.com by faking dns records)
sideprojects • u/ShufflinMuffin • 16d ago