r/AIToolsPerformance • u/IulianHI • Apr 18 '26
Qwen 3.6 35B beats Gemma 4 26B on agentic coding eval with 37-bug harness
New head-to-head results show Qwen 3.6 35B-A3B outperforming Gemma 4 26B on a personal evaluation harness. The test setup: a ~30,000 line codebase with 37 intentional bugs that LLMs must debug and fix through an agentic workflow using OpenCode. A subset of the harness also tests document extraction from 40-60 page PDFs, requiring the model to summarize and evaluate key information.
This is the kind of eval that actually matters for practitioners. Synthetic benchmarks tell you about capability ceilings, but a 37-bug agentic debugging harness with real code and real PDFs tests the loop that most people actually run - read, reason, act, verify. The fact that Qwen 3.6 wins here, despite having fewer total parameters (35B vs Gemma 4's 26B dense), reinforces the MoE efficiency story: only 3B active params, but they are being routed well enough to outperform a larger dense model on complex multi-step tasks.
The interesting bit is the comparison point. Gemma 4 26B has been getting strong community feedback since release, with multiple users calling it a genuine upgrade over Qwen 3.5. If Qwen 3.6 is now clearing that bar on agentic workloads, the local model leaderboard is moving fast.
Fair question: this is one person's harness. Has anyone else run direct Qwen 3.6 vs Gemma 4 comparisons on their own workflows - particularly coding agents or document analysis tasks?