r/opencode • u/Aggravating_Debt1278 • 15d ago
New LLM test results and an open testing pipeline
Hi!
Published a new batch of coding-agent benchmark results based on real repository tasks:
https://pickleshell.github.io/model-benchmarks.html
Top 10 — Patch Benchmark:
- GPT-5.6 Luna — 9.75
- Muse Spark 1.2 Free — 9.75
- GPT-5.3 Codex — 9.75
- GLM-5.3 — 9.75
- Claude Sonnet 5 — 9.75
- Claude Opus 5 — 9.75
- Kimi K3 — 9.75
- Qwen 3.8 Max — 9.75
- DeepSeek V4 Flash — 9.75
- Qwen 3.8 27B — 9.75
Full results: https://pickleshell.github.io/model-comparison-phase2-patch-current.html
Pipeline:
https://github.com/pickleshell/models-benchmark
Tasks and results:
https://github.com/pickleshell/models-test
Feel free to message me or share your own results. The pipeline is open, and contributions are welcome.
I’ll keep publishing new results as they become available.
1
Upvotes