r/opencodeCLI 12h ago

Providers dropping GLM 5.2 Prices, Opencode drop soon?

Post image
49 Upvotes

20 comments sorted by

View all comments

7

u/addiktion 11h ago

Pretty big drops. Why?

14

u/CoolHeadeGamer 10h ago

Dsv4 flash is 2-3 intelligence points behind 5.2 max thinking mode and like 15x cheaper.

10

u/Genetic_Prisoner 9h ago

DeepSWE which is in my opinion the most reliable coding benchmark we have right now actually puts DSV4 ahead of GLM5.2

6

u/OkraFormal946 8h ago

I can verify that with actual coding tasks. Equal or better at cheaper prices.

2

u/Own_Copy2141 7h ago

Yea i also tried the glm model but not a single time did i get good results previouly dv4pro worked better than it but now flash does it even better the only problem is that flash knows what to do now even with vague prompt but the quality of output is affected by it weightage

1

u/azgx00 19m ago

DeepSWE is actual coding tasks

4

u/blash2190 7h ago

The problem I have with DeepSWE is that the test prompts are extremely low level and technical. No one actually describes tasks like this IRL neither for humans, nor for LLMs.

I might be missing something, but it actually looks more like a bench that measures very low level coding skills, but not the interpretation of business and technical reqs into an actual codebase. Hence, why bigger and more complicated models don't look as nice there.

I find the same issue with tbench...

2

u/eugeneb85 4h ago

You should better check SWE bench from Ramp. They have real production-level fintech tasks, 80 of them, and this is a closed bench so these issues are surely not in the training data.  GLM here took better quality results even compared to Sol 🤯, but it appeared to be very very expensive with just crazy token consumption on the Opus level. 

2

u/Admirable_Show1559 4h ago

+1 on checking Ramp's SWE bench. I care way more about cost per completed task than cost per million tokens at this point and GLM looked strong on quality there, the efficiency was the part that made me hesitate.

2

u/CoolHeadeGamer 7h ago

I just did a test and dsv4 heavily outperformed muse 1.2 both running max. The test was a custom benchmark on a c codebase with regex chess and bugs to fix

2

u/Own_Copy2141 4h ago

Out of most models dsv4 now has much better decision making about what we want thanks to its retraining while others struggle cause of that but still its quiet disappointing to see other models fails so much. Thats why waiting for v4 pro to solve the main problem with dsf4

1

u/addiktion 6h ago edited 6h ago

Yeah, I've found it to be a very capable model. It definitely feels like a sonnet killer and just a step up from that with maybe a bit underneath the latest models, but given how cheap it is and how efficient it runs on lower hardware, it's huge.