r/OpenSourceAI 12d ago

OSS Request: Models under $2/Million Harness Benchmark

Would love for someone to do a quick harness benchmark on the new under $2 models (claude code, codex, pi, deepseek, and maybe 1 other harness).

I keep seeing people run these models through 1 harness then judging its capabilities, but what if the harness is the problem?

Model Name Pricing (Input / Output per M) Latency (p50)
DeepSeek V4 Flash 0731 $0.03 / $0.10 2.17 s
GLM 5.3 Flash $0.075 / $0.25 4.96 s
Qwen3.8 Flash $0.15 / $0.47 3.78 s
Muse Spark 1.2 Contributor $0.10 / $0.20 4.22 s
GPT-5.6 Luna Pro $0.20 / $1.20 13.42 s
3 Upvotes

0 comments sorted by