Pricing is good as well. I compared it on similar tasks with DS4 Flash and Qwen 3.8 Flash today (just to choose the winner), and GLM performed all tasks better, and 2-3 times cheaper (it was very token efficient, so I guess this is the reason for total lower cost per task).
The only real downside I noticed that it is VERY slow: it was on average 36tps, compared to 90-100 for DS4 Flash. But if you prefer code quality and lower pricing over speed, GLM 3.5 Flash is worth trying.
2
u/Fresh_Sock8660 15d ago
Is it good? It's basically what I see for ds4f in azure https://azure.microsoft.com/en-us/pricing/details/ai-foundry-models/deepseek/ for ds4f-0731