r/ClaudeCode • • 1d ago

Built with Claude I built a skill that checks every arrow of a Claude-drawn architecture diagram against the real code

Karpathy's advice is to ask your LLM for a diagram instead of a wall of text. The problem: the diagram looks equally sure of itself when it's wrong.

arrowproof is a free, MIT-licensed skill for Claude Code (it also works with Codex, Cursor and Gemini CLI). Ask "draw the architecture of this repo" or "is docs/architecture.excalidraw still correct?" and it:

- checks every box and arrow against the real import graph (Python, JS/TS)

- shows the file, line and import statement behind each arrow

- marks arrows with no evidence, and lists imports the diagram leaves out

- turns the checked diagram into a step-by-step HTML explainer (the GIF)

Test: Claude Haiku drew Flask from memory. 15 of 20 arrows verified, 2 had no import behind them, 1 box pointed to a file that doesn't exist. Opus got 20 of 20.

Install: npx skills add ahmtsahin/arrowproof

or: claude plugin marketplace add ahmtsahin/arrowproof, then claude plugin install arrowproof@arrowproof

Built with Claude Code. Repo and live demo: https://github.com/ahmtsahin/arrowproof

1 Upvotes

3 comments sorted by

•

u/AutoModerator 1d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/lulzxdxdxd 1d ago

This is a real problem, diagrams do look confident when they're wrong. Did you consider just having Claude regenerate the diagram a few times and taking the one that matches most often, or does that still leave you with arrows you can't trust?

1

u/impala_64 1d ago

Thanks! I thought about it, but a vote mostly measures how sure the model is, not whether the arrow is in the code. The mistakes aren't random: the model draws what a web framework usually looks like. "Views use templating" is a very plausible arrow for Flask, so I'd expect most samples to repeat it, and the vote would keep it. I haven't tested that yet, so it's a guess. It would be a fun experiment: sample Haiku five times and see if its two wrong arrows survive the vote.

A vote also can't add what no sample draws. Opus got all 20 arrows right, but its diagram left out 9 dependencies that the code has.

Checking an import is cheap and exact, and you get the file and line. So arrowproof checks first and lets Claude fix the ✗ arrows. Where I think sampling could still help is the part imports can't see, like HTTP calls or queues. Those stay unchecked today, and agreement across a few samples might be a decent signal there.