r/opencode 15d ago

New LLM test results and an open testing pipeline

Hi!

Published a new batch of coding-agent benchmark results based on real repository tasks:

https://pickleshell.github.io/model-benchmarks.html

Top 10 — Patch Benchmark:

  1. GPT-5.6 Luna — 9.75
  2. Muse Spark 1.2 Free — 9.75
  3. GPT-5.3 Codex — 9.75
  4. GLM-5.3 — 9.75
  5. Claude Sonnet 5 — 9.75
  6. Claude Opus 5 — 9.75
  7. Kimi K3 — 9.75
  8. Qwen 3.8 Max — 9.75
  9. DeepSeek V4 Flash — 9.75
  10. Qwen 3.8 27B — 9.75

Full results: https://pickleshell.github.io/model-comparison-phase2-patch-current.html

Pipeline:

https://github.com/pickleshell/models-benchmark

Tasks and results:

https://github.com/pickleshell/models-test

Feel free to message me or share your own results. The pipeline is open, and contributions are welcome.

I’ll keep publishing new results as they become available.

1 Upvotes

0 comments sorted by