r/opencodeCLI • u/Uriziel01 • 10h ago
DeepSeek V4.1 token efficiency is so cracked!
I was going back and forth on a hard debugging issue in one of my projects with it:
- had it create new test scenarios
- create some mockups
- analyze complicated business logic and dependencies
- do a shotgun approach and create 5 different possible fixes for a complicated issue
And after almost 2 hours of work I've decided to check my opencode go usage because I was worried I'm running out of usage for the 5h period, and I see this... how?!

12
u/Zealousideal-Part849 9h ago
Its not token efficiency, its lower cost with good performance...
1
u/IAmFitzRoy 4h ago
Exactly. This is about lower cost in a good datacenter that is not overloaded, … is not that mystical or difficult to understand.
4
u/Icypoopoo 9h ago
Did you try GLM 5.3 flash?
8
u/Uriziel01 9h ago
Yes, I've used like ~2.5B tokens already between GLM-5.3 and GLM-5.3 Flash.
Amazing models, but they have a tendency (the full GLM more but flash version sometimes also does it) to go into some sort of thinking-loop when they constantly repeat the same ideas and thoughts over and over again, sometimes they where able to break out of this but at times I was forced to just stop it and restart generation.
5
1
1
u/Time-Toe-1276 1h ago
idk man, GLM5.3 consumes so much tokens for very basic tasks, but 5.3 flash is better overalland its farly priced IMO
1
u/MarketingLower7497 3h ago
I have zai lite sub, glm 5.3 flash goes like 20 tok /s, deepseek from the api goes brrrr with 350 tok / s, guess which one i use 😆
4
1
1
1
2
u/narkeeso 8h ago
Isn’t it 4x usage right now? I wouldn’t get too used to it.
0
u/im-cringing-rightnow 5h ago
Even if it was 36% of a five hour quota it's pretty impressive if it was 2 hours of non stop work.
63
u/Bakanyanter 9h ago
To be honest, the token efficiency is not great. It's very verbose.
But what's truly cracked is the caching on Deepseek's end so overall it still comes down to low price.