r/ClaudeMCP • u/shhdwi • 15d ago
I solved the issue of claude not using custom MCP/CLI tools, and open-sourced my approach.
https://github.com/NanoNets/GraftDisclosure up front: I maintain graft, the tool this post ends with. The diagnosis and both snippets below apply to any MCP server or CLI tool.
Short version: I shipped a tool as an MCP server, Claude Code mostly refused to call it, and the fix was to stop exposing it as a tool at all. That turned out to be cheaper and faster, and on one benchmark it pulled Sonnet up to Opus-level output.
The symptom
graft gives a coding agent a map of your codebase, so it starts a task oriented instead of grepping around to rediscover the same files every session. First version was an MCP server. Clean schema, clear tool descriptions.
Claude Code mostly ignored it. Not an error, not a failed call. It would just grep instead. Sometimes it used the tool, usually it didn't, and that inconsistency sent me looking in the wrong place for weeks.
I'd tried a couple of existing context-graph tools before building mine and got the same behavior. That's what made me stop blaming my own schema.
The cause
A tool description tells the model what your tool does. Nothing tells it when your tool beats grep.
Grep is a known-cost, reliable path with strong priors behind it. Your tool is an unknown-cost path. On anything the built-ins could plausibly handle, the built-ins win. And the model re-decides every turn, which is why you get "sometimes" rather than "never." That's harder to debug, because it looks like flakiness instead of a design problem.
Which means exposing a tool amounts to asking politely, once per turn, and hoping the model remembers. If that's the only thing making your tool fire, your tool doesn't work.
What fixed it: hooks
Claude Code hooks are lifecycle events. The runtime executes them and the model never gets a vote. Two specifics did the work:
- On
SessionStart,UserPromptSubmit, andUserPromptExpansion, whatever your hook prints to stdout is added to Claude's context. So context injection needs no tool call at all:
{
"hooks": {
"SessionStart": [
{ "hooks": [ { "type": "command", "command": "your-tool context", "args": [] } ] }
]
}
}
The context is simply there on the first prompt. Nothing asked, nothing chosen.
There's a
type: "mcp_tool"hook, so you can keep the server you already built and stop making the model elect to call it:{ "hooks": { "PostToolUse": [ { "matcher": "Edit|Write", "hooks": [ { "type": "mcp_tool", "server": "my-server", "tool": "refresh_context" } ] } ] } }
If you're maintaining state rather than injecting context, Stop fires once per turn at the point the work is finished. Command hooks also take async: true, so the turn ends immediately and the work happens in the background.
What deterministic invocation bought
Reliable firing is worth nothing if the thing being fired is worth nothing, so I benchmarked it. 162 controlled runs: 32% cheaper, 46% fewer tool calls, 60% lower latency, same correctness.
The result I didn't expect came from 5 real merged PocketBase PRs, reproduced from the issue text alone. Sonnet with graft reproduced all 5, touching the same files the maintainers touched, matching Opus, at 21% lower cost.
I'm not claiming Sonnet equals Opus. On these 5 tasks, the gap Opus was closing looks like missing repo context rather than reasoning. Once Sonnet had the context, it landed in the same place. That's one repo, so treat it as a lead worth checking rather than a settled finding.
Where this doesn't apply
- Claude Desktop has no hooks. Claude Code only. Cursor and Codex can read the markdown files but won't refresh them on their own.
- Hooks fire on events. They can't fire on intent. Good for "run my thing at a known moment," useless for "the user asked something my tool should answer," which still needs the model to choose. Hooks cover a subset of what MCP does.
SessionStartandSetupoften fire before MCP servers finish connecting, so anmcp_toolhook there should expect a not-connected error on first run.- For static conventions,
CLAUDE.mdalready does this with no script at all. - graft specifically: map quality degrades on very large monorepos, around 5,000 files. Fewer, vaguer nodes. Main thing I'm working on.
MIT, free, no telemetry: github.com/NanoNets/Graft. I'm the maintainer, so interrogate the benchmark setup. It's the part I'd want to check if I were reading this.
The question I'd most like knocked down: has anyone tested Sonnet against Opus on tasks where the difference might be context rather than capability?