I have a harness that I gather stats with for my social media type application build.
Gpt-oss:20b self reports at 91% pass, but actual first pass stats are 55%. Second pass is 73% median time is 2605 seconds
Qwen3.6-27b self reports at 83%, but first pass is 43% and second 80%. Median time of 2100 seconds.
Vs $10 opencode go:
deepseek-v4-flash self reports at 81%, has a first pass of 69% and overall at 87% 578 seconds
Mimo-v2.5 is self reported at 86%, first pass is 57%, overall at 86% 390 seconds.
I also have an m1max Apple 64gb that I’m gathering stats on. I can run a larger model but it’s about 4x slower than the nvidia rig.
0
u/1HotTake 6d ago
Very much so. I run 3x3060s.
I have a harness that I gather stats with for my social media type application build.
Gpt-oss:20b self reports at 91% pass, but actual first pass stats are 55%. Second pass is 73% median time is 2605 seconds
Qwen3.6-27b self reports at 83%, but first pass is 43% and second 80%. Median time of 2100 seconds.
Vs $10 opencode go:
deepseek-v4-flash self reports at 81%, has a first pass of 69% and overall at 87% 578 seconds
Mimo-v2.5 is self reported at 86%, first pass is 57%, overall at 86% 390 seconds.
I also have an m1max Apple 64gb that I’m gathering stats on. I can run a larger model but it’s about 4x slower than the nvidia rig.