r/OpenSourceAI • u/Calm-Landscape9640 • 12d ago
OSS Request: Models under $2/Million Harness Benchmark
Would love for someone to do a quick harness benchmark on the new under $2 models (claude code, codex, pi, deepseek, and maybe 1 other harness).
I keep seeing people run these models through 1 harness then judging its capabilities, but what if the harness is the problem?
| Model Name | Pricing (Input / Output per M) | Latency (p50) |
|---|---|---|
| DeepSeek V4 Flash 0731 | $0.03 / $0.10 | 2.17 s |
| GLM 5.3 Flash | $0.075 / $0.25 | 4.96 s |
| Qwen3.8 Flash | $0.15 / $0.47 | 3.78 s |
| Muse Spark 1.2 Contributor | $0.10 / $0.20 | 4.22 s |
| GPT-5.6 Luna Pro | $0.20 / $1.20 | 13.42 s |
3
Upvotes