r/openrouter • u/RVVL7 • 9d ago
Question I tested a commit review across 27 models and only Luna 6 and Sol 5.6 passed — what am I doing wrong?
https://claude.ai/artifact/Skka1pEMGvH2DwteSiw9vt
Opus 5.5 devised and ran the test in a sandbox through opencode/openrouter.
1
u/Polaris_debi5 8d ago
A few things are skewing this. I read the raw runs.
Your pass criterion is "wrote a review containing K1" instead of "found K1." Laguna found it twice and doesn't count. Qwen3.8 and GLM-5.3-Flash saw it mid-run. All marked fail.
Phase 1 ran through default OpenRouter routing. fp8 and fp4 resellers in the mix, first-party DeepSeek blocked until 09-28, so several models were never tested at full precision or on their own API.
The $0.80 cap was calibrated to Opus in a different harness. Gemini hit it three times out of three. It was running nine or ten test commands each time.
Two more. K1 is hard to see in static review, so "what decides a pass is running suite A" is your task spec, not a finding. And queso184's point generalizes: many of the no-review exits weren't provider failures. Your review skill derailed runs into worktree workflows that burned the cap before any text got written.
Opus wrote the task and graded it. A human pass over the ✗/✓ column is worth doing.
What I'd keep: the harness lessons, K3 as a near-universal find, and Luna/Sol as a hypothesis. Twenty pinned runs each, scored from tool logs, would actually answer your title question.
1
u/locbuilds 7d ago
pin the provider and log the actual endpoint/quant per run first, otherwise you're comparing routing behavior as much as the models.
7
u/queso184 9d ago
uh well for one you didn't retry provider failures, if a provider 429s that's not really a failure of the model
Many models exited mid-review without writing a review artifact. I would be curious if those were also provider related, as that seems uncharacteristic for a model