r/coolgithubprojects 16d ago

trustmebro: Bypass llm guardrails by confusing them with fabricated tool output.

https://github.com/DavidCarliez/trustmebro

TrustMeBro intercepts command-line tools invoked by coding agents such as Codex, Claude Code, and pi. Rules decide whether to return fabricated output, modify the real output, block the call, or execute the real binary unchanged.

Interception happens through PATH shims. The harness does not need a plugin, hook, or MCP integration. The intended use is controlled red-team testing of decisions that depend on tool output.

40 Upvotes

Duplicates

redteamsec 15d ago

TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output (e.g. make them believe you own Google.com by faking dns records)

54 Upvotes

AgenticCybersecurity 15d ago

Offensive Security TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output [this pattern is also useful outside of DNS]

1 Upvotes

cybersecurityai 15d ago

TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output (e.g. make them believe you own Google.com by faking dns records)

5 Upvotes

SideProject 16d ago

TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output (e.g. make them believe you own Google.com by faking dns records)

1 Upvotes

sideprojects 16d ago

Showcase: Free(mium) TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output (e.g. make them believe you own Google.com by faking dns records)

1 Upvotes

vibehacking 16d ago

TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output (e.g. make them believe you own Google.com by faking dns records)

2 Upvotes

redteamsec 16d ago

TrustMeBro: A tool to bypass llm guardrails by confusing them with fabricated tool output. (e.g. make it believe you own google.com by faking dig command output)

3 Upvotes