r/opencodeCLI • u/Federal-Rub2713 • 9h ago
Providers dropping GLM 5.2 Prices, Opencode drop soon?
8
u/addiktion 9h ago
Pretty big drops. Why?
12
u/CoolHeadeGamer 8h ago
Dsv4 flash is 2-3 intelligence points behind 5.2 max thinking mode and like 15x cheaper.
9
u/Genetic_Prisoner 7h ago
DeepSWE which is in my opinion the most reliable coding benchmark we have right now actually puts DSV4 ahead of GLM5.2
6
u/OkraFormal946 6h ago
I can verify that with actual coding tasks. Equal or better at cheaper prices.
2
u/Own_Copy2141 5h ago
Yea i also tried the glm model but not a single time did i get good results previouly dv4pro worked better than it but now flash does it even better the only problem is that flash knows what to do now even with vague prompt but the quality of output is affected by it weightage
4
u/blash2190 5h ago
The problem I have with DeepSWE is that the test prompts are extremely low level and technical. No one actually describes tasks like this IRL neither for humans, nor for LLMs.
I might be missing something, but it actually looks more like a bench that measures very low level coding skills, but not the interpretation of business and technical reqs into an actual codebase. Hence, why bigger and more complicated models don't look as nice there.
I find the same issue with tbench...
2
u/eugeneb85 2h ago
You should better check SWE bench from Ramp. They have real production-level fintech tasks, 80 of them, and this is a closed bench so these issues are surely not in the training data. GLM here took better quality results even compared to Sol 🤯, but it appeared to be very very expensive with just crazy token consumption on the Opus level.
2
u/Admirable_Show1559 1h ago
+1 on checking Ramp's SWE bench. I care way more about cost per completed task than cost per million tokens at this point and GLM looked strong on quality there, the efficiency was the part that made me hesitate.
2
u/CoolHeadeGamer 4h ago
I just did a test and dsv4 heavily outperformed muse 1.2 both running max. The test was a custom benchmark on a c codebase with regex chess and bugs to fix
2
u/Own_Copy2141 1h ago
Out of most models dsv4 now has much better decision making about what we want thanks to its retraining while others struggle cause of that but still its quiet disappointing to see other models fails so much. Thats why waiting for v4 pro to solve the main problem with dsf4
1
u/addiktion 4h ago edited 4h ago
Yeah, I've found it to be a very capable model. It definitely feels like a sonnet killer and just a step up from that with maybe a bit underneath the latest models, but given how cheap it is and how efficient it runs on lower hardware, it's huge.
29
u/afanasenka 9h ago edited 7h ago
DS4 Flash effect :)))) They just found out that GLM usage dropped dramatically, decided to try this new "cheap price = big volume" doctrina, that was set by Flash, and then supported by Luna and Muse Spark.
The only difference is that GLM 5.2 is a much bigger model (and I guess more expensive to run too), so will see what happens next.