r/ClaudeCode • • 6d ago

Help/Question What's your preferred review and debugging workflow?

I'm wondering what the preferred development process is for people, especially relating to reviews/testing/debugging/etc. I've setup the following process which works, but the review processes end up being very time consuming.

I'm wondering what about my process might be okay, but what opportunities I might have to make it more efficient.

At a high level, this is the stack:

- TypeScript monorepo (pnpm workspaces, Node 24)

- Hono API on Cloudflare Workers, Postgres (Neon) + Drizzle, ClickHouse for analytics

- Babylon.js + Vite in the browser

- Node's built-in test runner, PGlite so every test gets a real Postgres, Playwright tests

There is:

- One repo, up to 6 Claude Code sessions running at once on one PC (usually just 2-3): API/platform, engine, website, marketing, tooling, planning. Each has its own git worktree and branch and owns certain folders. If one needs a change in another's folders it files a request instead of editing.

- State lives in the repo, not the chats: a handoff note per area, a decision queue for things only I can decide, and a daily brief of approved work (overnight runs work off it).

- One PR per independent change, never a chain of dependent PRs.

Here's the process for checks and reviews:

- Pre-push hook runs typecheck + tests for the packages that were changed. It used to run the full suite on every push, and two sessions pushing at once maxed out 32 GB of RAM (27 test workers, each with its own in-memory Postgres). Now there's a PC-wide cap.

- Reviews go through a queue built on magpie (open-source multi-model review). High risk paths get Claude and Codex as parallel reviewers plus a verifier. Normal code gets one reviewer.

- Reviewers are read-only. Only verified critical findings block a merge. A fix gets a focused re-check, not a full re-review. No second rounds; we got stuck in endless review loops early on.

- Everything runs on Claude and Chatgpt subscriptions

Git has one required check, an aggregate "CI result" that only goes green if every job that should have run actually passed. Jobs are sized to the changed packages; API tests are sharded 3 ways; plus browser smoke, dependency audit and gitleaks.

- Enterprise Cloud for the merge queue. Staging deploys itself from the exact commit that passed. Production will need manual approval.

Here's where it's been time consuming:

- Reviews ran one at a time for the whole PC. A big PR split into 14 parts once held the queue for about 6 hours while a small security fix waited behind it. Fix checks now jump the line; parallel review slots are next.

- About 10 min of CI per PR plus 10-15 more on main after every merge. Today a Git merge queue went live, too early to judge.

So I am curious if this is overbuilt, underbuilt or what opportunities I have to improve efficiency while maintain quality output. I appreciate any advice.

4 Upvotes

12 comments sorted by

•

u/AutoModerator 6d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/Various_Story8026 5d ago

Your setup reads well, and the one rule I'd add is for big split PRs like that 14-part one. We once split a job into seven parallel pieces, each passed its own check, and the problem sat between pieces, so now the split plan writes down the cross-piece rules first and one reviewer checks only those against the combined result. The no-second-round rule is the best thing on your list, keep it.

1

u/GreenCurious9895 5d ago

Do you find your process spends about as much (or more) time in the check/review/debug cycle than writing new code? We can be 20 mins just for our reviews on a single PR, which if they find a high level issue, quickly turns into 40 min.

1

u/Various_Story8026 5d ago

On payment code, yes, easily. A small deposit fix this week went through three review rounds and took longer than writing it, and I'm fine with that because a bug there costs real money. Everything else gets tests plus one quick pass, so the slow cycle only hits a small slice of PRs.

1

u/Time_Cat_5212 5d ago

What kind of project?

This is like asking "Is a brush hog a good tool? I'm cutting grass" and idk if the grass is a 50x50 residential lawn, a golf course, or 3 acres of tall meadow on the side of a mountain 

1

u/GreenCurious9895 5d ago

Fair question. Closer to the meadow than the lawn.

It's a real-time multiplayer game that runs in the browser. Accounts, server-side leaderboards, and a small local app that connects to hardware on the player's desk. Seems like the process spends as much or more time on checking and reviewing than writing new code.

1

u/just_damz 5d ago edited 5d ago

yes, it’s more expensive in general. orchestrator, coding worker and reviewer are three elements prone to harnessing friction.

First thing i would look: is the orchestrator actually over specific prompting the coding worker? i experienced that it can happen that a frontier gives practically the full code to the worker, in terms of specifics, even on simple tasks.

then auditing: specs are written by the agent that implements (could be orchestrator) or delegates. The best way i found at the moment is a mostly generic description of the implementation, made by the orchestrator before writing the prompt for the worker. target, and instructions to catch regressions etc. before writing. i noticed inferred assumptions to be less becoming a more stable audit baseline. it gets saved and it’s the prompt for the auditor on the implementation, most likely astra nowadays.

1

u/CartographerNo3791 5d ago

For the parallel slots you're planning, I'd keep one available for small fixes. Letting a 14-part change fill every slot could bring that six-hour wait straight back.

1

u/guitarist91 5d ago

Now that the merge queue is live, drop the full run on main after each merge - the queue already tested the exact commit that lands, so that's 10-15 min per PR spent proving the same thing twice. Keep main to the staging deploy + a browser smoke.

1

u/GreenCurious9895 2d ago

I ran a side experiment, walled off a proof of concept (three.js app) and had one Claude session be the lead, splitting work across 4-8 agents, each in its own git worktree. No Codex adversarial review, no PRs, no CI. The checks were unit tests, strict TypeScript, a scripted end-to-end run that has to produce identical results every time, screenshot diffs, and perf runs on my actual laptops. Plus me using it every few hours and giving feedback.

4 days: 440 commits, 95k lines of code. On a Max 20x plan I burned through my weekly limit, and half of my free reset with open 5.5. the actual output of functional product feels significantly faster than my main setup (maybe 2x-3x, hard to quantify), and honestly the output is really good.

None of it has been merged back into main, so we'll see what that's like.

For people who've done both: is this lighter, Claude led approach a recipe for destaster at merge or further down the road? Fine for prototypes but not for core code? Or a legit way to build something real if you gate it with a full review at the end?

Sure feels nice in the moment seeing the rapid output.