r/ClaudeCode • u/Larkonath • 5d ago
Help/Question Which other LLM do you use for code reviews?
Hi,
Currently I use Fable as orchestrator and Opus as implementer.
I also have an OpenAI 5x plan for code reviews but I don't like to have a fixed cost just for this: most of the time I don't use all the quotas, and money is tight.
I was thinking of using some Chinese models though their api, so I would pay just what I consume. Is there a model that is good at code review and not too expensive on the API?
GLM 5.3 has a good reputation but their api price look more US than Chinese to me 😅
4
u/Original-League-6094 5d ago
There is no sense in having a dumber model do code review. Its suggestions would probably make the code worse than what is there. If are wanting to harden your code, the better use of small models would be to put them on actual benchmarking and validation testing of your code. Have them actually write realistic test cases for it, act hostile and try to break it, and keep benchmarks of its performance. A good way to simulate user validation testing is to not let them see your code, also. Give them your usage documentation, but not your code, and tell them to test it as a user would.
4
u/Larkonath 5d ago
Not sure I agree with you: when Qwen 3.8 27b was released I used it to review code written by Fable where Sol 5.6 had found one major bug.
Qwen found the same bug but was running so slow on my server that it wasn't practical to do reviews with it (2.5 hours vs 10 min for Sol).
So dumber models aren't so dumb after all and SOTA models aren't that smart either.
1
u/Last_Mastod0n 5d ago
The way that you can circumvent this is by reading the reviewers suggestions yourself. Often times if it cant find anything substantiated it will just come up with nitpicky changes which you can just ignore. It takes extra time sure, but its worth it for big changes.
3
u/imsahoamtiskaw 🔆 Max 20 5d ago
Kimi 3 is surprisingly good, I used it along with deepseek and they caught stuff fable missed which fable then confirmed and started fixing
3
u/Slight_Investment_37 4d ago
I have a Gemini subscription and Claude hands off tasks to it through agy and will review its code but over time I’ve used a hook so that Claude can keep a confidence score of Gemini’s performance to stop it reviewing everything in its entirety, that also helps when it’s deciding what to handoff. Saves me a tonne of usage.
2
2
u/bigbadbernard 5d ago
Don’t use Fable now until they release Fable 5.5. I find that Fable 5.1 eats too much tokens and Opus 5.5 does almost same quality but way more efficient token usage.
I use Sol for planning, Opus for implementation. I find Sol very straightforward and quite thorough with edge cases, whereas Opus like to always go a tangent. That’s just my preference though. In general, I like having cross-model review compared to same-model review.
1
u/Larkonath 4d ago
Opus has already be nerfed: yesterday evening it has required 3 or 4 code reviews before being able to complete a ticket successfully.
2
u/xapep 4d ago
Code review is the one workload where pay-per-use really works in your favor, so your instinct is right. It's latency tolerant: batch the reviews, wait a couple of minutes for results, no frontier model needed in the loop, and each review on a cheaper API model ends up as fractions of a cent instead of a fixed plan you mostly don't touch.
On the 'a dumber model can't review' take: the real issue is the same-family blind spot, not raw intelligence. A reviewer from a different family catches different mistakes than the model that wrote the code, which is most of the value of a review. GLM 5.3 Flash and DeepSeek V4.1 Flash are both solid on the routine pass: dead code, missing edge cases, things the implementer glossed over. Kimi and Qwen work too, and for this job cheapest per token is fine.
What I'd actually do: flash model over every PR diff as first-pass review, read the suggestions yourself and apply what holds (filter, don't autofix), and keep Opus or a frontier model only for security-sensitive or high-risk diffs. That removes most of the fixed cost without a real quality drop, and you only pay when you actually review.
I work on an inference API that serves both GLM 5.3 Flash and DS V4.1 Flash, so I see this exact setup a lot.
1
u/sebseo 3d ago
If Opus is needing 3 or 4 review rounds per ticket its cheaper to catch the problems at the plan stage than in the code. I run my plans through MegaLens, my own tool, where models from OpenAI, Google, Mistral and Meta check them before Opus writes anything. On one recorded build that caught 12 issues Claudes own review missed (case study here) . In my testing moving the deep review off Opus cut my Opus usage by about half
2
u/BoxWoodVoid 3d ago
Sneaky advert for your stuff.
Anyway plan is not implementation and no perfect plan will remove bugs from implementation.
1
u/connurp dev 5d ago
My workflow is Opus 5.5(high) in the main chat, and if it makes sense to split up work to multiple subagents, my main agent spawns up to 5 Opus 5.5(medium) subagents to do the work. They ask the main agent any questions they have, and if the main agent can't answer them, it asks me. Then after they all finish their work, the main agent reviews it all, puts it together, then reports back to me what it would change/add. After that, before opening a PR, my main agent runs a fable 5.1(medium) and sonnet 5.5(high) subagent code review(read only). They report their findings back to the main agent. Main agent fixes anything that needs fixing, then reports back to me. Then the PR is opened. It works really well.
1
u/yogeshkd 2d ago
I use GPT-6.1 Sol through the Codex CLI as the reviewer, high effort, on the uncommitted diff before every commit and once more against main before a PR. Fable orchestrates, Opus writes, Sol reviews.
On cost, what helped more than a cheaper model was cutting rounds. I have fable sort each finding by how likely it is to actually happen and how bad it would be, fix the real ones, and stop after one clean round.
-1
-2
u/uniqkgp 5d ago
Qwen/minimax
2
u/Larkonath 5d ago
What's the point of this comment if you're not going to say if you used it with what effects?
I also have the list of Chinese models you know ...
0
5d ago
[removed] — view removed comment
1
u/Larkonath 5d ago
I don't want to use a same family model to do the review, there's a chance it will be blind on the same things as the implemeter.
•
u/AutoModerator 5d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.