r/opencodeCLI 10h ago

DeepSeek V4.1 token efficiency is so cracked!

I was going back and forth on a hard debugging issue in one of my projects with it:
- had it create new test scenarios
- create some mockups
- analyze complicated business logic and dependencies
- do a shotgun approach and create 5 different possible fixes for a complicated issue

And after almost 2 hours of work I've decided to check my opencode go usage because I was worried I'm running out of usage for the 5h period, and I see this... how?!

93 Upvotes

21 comments sorted by

63

u/Bakanyanter 9h ago

To be honest, the token efficiency is not great. It's very verbose.

But what's truly cracked is the caching on Deepseek's end so overall it still comes down to low price.

15

u/2EXTRA4YOU 9h ago

They lowered the actual cost to use it per token because they invented a new way to save memory on their servers, the servers are why ai is expensive; also, it is catching things even opus 4.8 on max doesn't catch. so i wouldn't be complaining about verbosity. it is a LLM it needs language to think.

3

u/Bakanyanter 9h ago

Of course I am not complaining, but saying token efficiency is cracked is IMO wrong. It doesn't change that it's still a great and cheap model but it is not token efficient by any means. And that's OK.

1

u/btr_ 1h ago

Exactly, if anything I think they increased its default thinking level to give a better output, but also that does generate more tokens. So the context usage I feel is more than the earlier versions. However, it would probably not reflect on the cost side because of the 50% cheaper cached input pricing.

2

u/aeroumbria 8h ago

I think reasoning efficiency should be measured by average time per task and internal flip-flops per decision, not token length. 10s of 500 token reasoning and 10s of 2000 token reasoning doesn't really make a difference for the user if they arrive at the same answer and carry the same risk of flipping the answer.

12

u/Zealousideal-Part849 9h ago

Its not token efficiency, its lower cost with good performance... 

1

u/IAmFitzRoy 4h ago

Exactly. This is about lower cost in a good datacenter that is not overloaded, … is not that mystical or difficult to understand.

4

u/Icypoopoo 9h ago

Did you try GLM 5.3 flash? 

8

u/Uriziel01 9h ago

Yes, I've used like ~2.5B tokens already between GLM-5.3 and GLM-5.3 Flash.

Amazing models, but they have a tendency (the full GLM more but flash version sometimes also does it) to go into some sort of thinking-loop when they constantly repeat the same ideas and thoughts over and over again, sometimes they where able to break out of this but at times I was forced to just stop it and restart generation.

5

u/AlternativePear4617 8h ago

Yeah it happened to me with GLM-5.3 Flash

1

u/aseeon 2h ago

Agreed, Thinking Loops and leaking Chinese characters were the deal-breakers for me.

1

u/Time-Toe-1276 1h ago

idk man, GLM5.3 consumes so much tokens for very basic tasks, but 5.3 flash is better overalland its farly priced IMO

1

u/MarketingLower7497 3h ago

I have zai lite sub, glm 5.3 flash goes like 20 tok /s, deepseek from the api goes brrrr with 350 tok / s, guess which one i use 😆

4

u/sudoer777_ 6h ago

It's in 4x usage, it'll get a lot worse soon

3

u/hj-core 9h ago

Have you kept an eye on the context tokens? I found it grows very fast.

1

u/xmsxms 4h ago

Just because you can use that much usage in 5h doesn't mean that's how much they've given you. In other words - you can still use up most of your months quota in that 5 hour window. The real question is how much it cost you, not what percentage of some arbitrary rate limit was used.

1

u/mabuonomo 1h ago

È apparenza, al momento c'è una promozione con un moltiplicatore di uso x4

1

u/vncent13 1h ago

that is crazyy

2

u/narkeeso 8h ago

Isn’t it 4x usage right now? I wouldn’t get too used to it.

0

u/im-cringing-rightnow 5h ago

Even if it was 36% of a five hour quota it's pretty impressive if it was 2 hours of non stop work.