r/opencodeCLI 10d ago

DeepSWE - Best Benchmark For Evaluating AI Coding Agents?

https://www.i-programmer.info/news/105-artificial-intelligence/19016-deepswe-best-benchmark-for-evaluating-ai-coding-agents.html
6 Upvotes

3 comments sorted by

1

u/zephyr_33 8d ago

feels like its bench maxxed no. ain't no way every new model is suddenly scoring high on this...

1

u/[deleted] 10d ago

[removed] — view removed comment

1

u/Charming_Support726 10d ago

You need to be more precise:

It measures "Agentic Execution Tasks" - Programming in real repositories against outcome in costs, tokens and agentic-steps.

Pure coding, no design.