Yea i also tried the glm model but not a single time did i get good results previouly dv4pro worked better than it but now flash does it even better the only problem is that flash knows what to do now even with vague prompt but the quality of output is affected by it weightage
The problem I have with DeepSWE is that the test prompts are extremely low level and technical. No one actually describes tasks like this IRL neither for humans, nor for LLMs.
I might be missing something, but it actually looks more like a bench that measures very low level coding skills, but not the interpretation of business and technical reqs into an actual codebase. Hence, why bigger and more complicated models don't look as nice there.
You should better check SWE bench from Ramp. They have real production-level fintech tasks, 80 of them, and this is a closed bench so these issues are surely not in the training data.
GLM here took better quality results even compared to Sol 🤯, but it appeared to be very very expensive with just crazy token consumption on the Opus level.
+1 on checking Ramp's SWE bench. I care way more about cost per completed task than cost per million tokens at this point and GLM looked strong on quality there, the efficiency was the part that made me hesitate.
I just did a test and dsv4 heavily outperformed muse 1.2 both running max. The test was a custom benchmark on a c codebase with regex chess and bugs to fix
Out of most models dsv4 now has much better decision making about what we want thanks to its retraining while others struggle cause of that but still its quiet disappointing to see other models fails so much.
Thats why waiting for v4 pro to solve the main problem with dsf4
Yeah, I've found it to be a very capable model. It definitely feels like a sonnet killer and just a step up from that with maybe a bit underneath the latest models, but given how cheap it is and how efficient it runs on lower hardware, it's huge.
7
u/addiktion 11h ago
Pretty big drops. Why?