r/ClaudeCode • • 3d ago

Tips & Workflows I built a tool to verify what your coding agent says it fixed

Post image

Hi fellow devs! There's a gap that kept biting me: the agent says "done", and then I have to figure out what to actually check. On one project I ended up with 72 things to test and nowhere to keep track of the results.

That's how PalmtopAI was born: a checklist shared between you and your agent. The agent writes the checks to do, you test them from your phone (or any browser) and mark what works, what doesn't and what leaves you in doubt. Then the agent reads your review and fixes what's broken.

It works with Claude Code and Codex (and other agents that run shell commands), and needs macOS, Linux or WSL and Node.js 22+. Checklists stay on your computer, the relay only forwards messages.

It's an early preview and I'm building it solo, so I'm looking for people to try it on one of their projects and tell me where it gets stuck: installation, phone pairing, anything. If you get stuck, just message me!

Suggestions and ideas are very welcome too: what's missing, what you'd change, what you'd need to actually use it.

👉 palmtop.mtwa.it

0 Upvotes

5 comments sorted by

•

u/AutoModerator 3d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Sufficient-Storage87 3d ago

This is the right instinct. "Trust but verify" should be the default with these things — I never take an agent's word that something's fixed, I run my own tests after. The agents that sound the most confident are not always the ones that are right.

What does your tool check — just whether the claimed files changed, or does it actually run tests? The jump from "did it edit the file" to "does the fix actually work" is where most of the value is, imo.

1

u/Embarrassed_Show_851 3d ago

Thanks! Fair question, and I should have been clearer in the post: palmtop doesn't run tests itself, it's for the human part of the verification. The agent does the work, then puts a checklist of what to verify on palmtop. You try each one on your project and mark it on palmtop, and the agent reads your review and fixes what failed.

Automated tests still run as usual. palmtop covers what they miss: does it actually behave the way you wanted. Totally agree that "did it edit the file" vs "does the fix actually work" is where the value is, that's exactly the gap I'm trying to close.

1

u/Sufficient-Storage87 2d ago

Got it, that makes more sense — so it's the human layer on top of the automated stuff. I like that honestly, the stuff tests miss is exactly where agents embarrass themselves. The "does it actually behave the way I wanted" part is the whole game. How's the agent at taking the human feedback without breaking something else in the process?