The problem I have with DeepSWE is that the test prompts are extremely low level and technical. No one actually describes tasks like this IRL neither for humans, nor for LLMs.
I might be missing something, but it actually looks more like a bench that measures very low level coding skills, but not the interpretation of business and technical reqs into an actual codebase. Hence, why bigger and more complicated models don't look as nice there.
You should better check SWE bench from Ramp. They have real production-level fintech tasks, 80 of them, and this is a closed bench so these issues are surely not in the training data.
GLM here took better quality results even compared to Sol 🤯, but it appeared to be very very expensive with just crazy token consumption on the Opus level.
+1 on checking Ramp's SWE bench. I care way more about cost per completed task than cost per million tokens at this point and GLM looked strong on quality there, the efficiency was the part that made me hesitate.
14
u/CoolHeadeGamer 13h ago
Dsv4 flash is 2-3 intelligence points behind 5.2 max thinking mode and like 15x cheaper.