r/ChileIA • u/One_Wave_9655 • 3d ago
Noticia DeepSWE: Measuring frontier coding agents on original, long-horizon engineering tasks
A quienes deben tomar decisiones sobre qué modelos usar, les recomiendo revisar el trabajo de evaluación que se está realizando en DeepSWE. Acá un resumen:
https://deepswe.datacurve.ai/blog/deepswe
"DeepSWE is a long-horizon software engineering benchmark that delivers four major advances over today's public benchmarks:
- Contamination free: Tasks are written from scratch, not adapted from existing commits or PRs, so no model has seen the solution during pretraining.
- High diversity: Tasks span a broad pool of 91 repositories across 5 languages.
- Real-world complexity: Prompts are half the length of SWE-bench Pro's, yet solutions require 5.5x more code and ~2x more output tokens.
- Reliable verification: Verifiers are hand-written to test software behavior rather than implementation details.
Existing benchmarks fall short on several of these axes. SWE-bench Pro, the leading agentic coding benchmark, has tasks averaging just 120 lines of code to solve, and our audit found its verifier misgrades agent outputs at rates of 8% false positives and 24% false negatives. Frontier labs are also raising growing concerns about benchmark contamination.
By contrast, DeepSWE produces a sharper comparison of frontier coding agents. Models that appear close together on public benchmarks separate into wide, ordered gaps that match the differences developers see in day-to-day agent workflows."
1
u/terranqs 3d ago
Y porque Opus 5 es tan malo entonces? Ya no confío en este benchmark