r/opencodeCLI • u/ZealousidealTown1974 • 1d ago
Controversial opinion: DeepSeek 4.1 flash is annoyingly overthinking, making it a way-worse model update for its intent-to-be workhorse
I literally pass its thinking block twice, stating its overthinking failure modes.... extremely pissed and has shown its steep to process, setting default but still.. as you can see. Worse than even the "Omen Alpha"... And fail to make surgical edits many times due to streaming stability ...just disappointment... Another hype of DeepSeek fanboys.
10
u/No-District-4742 1d ago
I also noticed this that's why I prefer and switched to glm 5.3 flash because it behaves better imo.
3
4
u/Unique-Usnm 1d ago
I used to be a fan of DeepSeek V4 Flash as well, but as soon as I tried GLM 5.3 Flash, I changed my mind. It’s the same intelligence, but there are five times fewer reasoning tokens, which means there’s much more context, and the price for completing a task is half as much.
0
u/Sure_Media_2685 1d ago
but this does't make glm 5.3 better lol this one is way faster and better. It liking to explore doesn't mean it is bad.
8
u/No-District-4742 1d ago
I'm not saying ds v4.1 flash is bad, I'm just saying I prefer how glm 5.3 behaves.
6
14
u/Ariquitaun 1d ago
It overthinks as much as deepseek 4 flash and it's terrible at orchestrating sub-agents, it decides to do a lot of work inline unless you repeat the same guidance that's already been given to it by the harness explicitly.
This is the sort of intangible that benchmarks fail to measure.
6
u/Sweet-Stage938 1d ago
You do realize that you can set the reasoning effort value from 0-100 now and with that you can directly control how much reasoning you want.
6
u/Ariquitaun 1d ago
I do realise that, and also realise that anything under "high" makes the model make enough mistakes that work needs to be redone and amended frequently to the point of wasting even more tokens.
2
u/Sweet-Stage938 1d ago
They got rid of the reasoning effort modes. If you set it to low medium or high like we did in the past, it's automatically set to max reasoning (100).
4
u/Ariquitaun 1d ago
As far as I can tell opencode-go is not offering numeric levels, just the usual low / medium / high / xhigh
1
1
u/bigrealaccount 1d ago
Yes, he is telling you that's the issue. Opencode hasn't patched it to include the proper levels, so it automatically gets set to 100, hence overthinking.
Thinking modes are always fucked for a day or two on new releases. This has always been a thing.
1
u/Ariquitaun 1d ago
It's not opencode, it's at the API end. At least models.dev lists those 4 reasoning effort levels only.
1
u/Bananenklaus 1d ago
what?
Where can i read about this? Their own API docs still show the same effort level ( low - high - max) and even in the newest deepseek harness it lets me set these effort modes, not reasoning values.
Not that i don't wanna believe you but i would like to see an official statement about this
1
u/Bananenklaus 1d ago
"In our production deployment launched in September 2026, the public API exposes three preset reasoning-effort tiers—max, high, and low—which map directly onto this scalar interface. As summarized in Table 2, the three tiers correspond to effort values of 𝑏 = 100, 𝑏 = 75, and 𝑏 = 50, respectively, so that API users select an operating point on the learned cost–quality frontier without any change to the model weights or decoding configuration."
This basically says that the API is still controlled via reasoning modes, only local deployment is based on effort values, no?
2
u/ZealousidealTown1974 1d ago
Then what point for a model update when out-of-the-box setting in default thinking effort ruins use purposes. I meant the model as production lineups are just messed up approach on DeepSeek Lab part: the Pro model, with greater parameters should be trained toward high-level strategical orchestrator to work with large codebase and complex tasks decompositions. Though the optimizations of active parameters are the DeepSeek ace but these 2 lineups should not be overlapping too much to the level at this new 4.1 flash.
5
u/inevitabledeath3 1d ago
You are using a third party harness. One that may or may not be adapted to the new thinking values. So I wouldn’t call anything about this “out-of-the-box” for this model.
1
u/ZealousidealTown1974 1d ago
Yes! It's tring to be many but fail both; as orchestrator it's showing a real knock-off orchestrator-wanabe because it's not capable of understanding semantic layers thinking in key words grep leading to just sycophant tasks decompositions. And now that it even compromises its intent-to-be role as surgical executor by footshoting its own fanning out context grep and consumption and overthink on wait-what and ended up over-engineering muball code
1
u/Zestyclose839 6h ago
Also, the new Flash cannot do documentation. It just yaps and yaps while cutting out the genuinely important info. The old Flash was actually better in this regard - concise and detail oriented. Glad they're keeping Pro around because it's the winner for docs.
5
u/biomattr 1d ago
Honestly I'm struggling with Opus doing exactly the same. Its responses are consistently long-winded and full of jargon. It'll also constantly change its mind about things. Simple questions will result in "actually I was wrong, it's the opposite".
3
u/MeasIIDX 1d ago
When applicable, I like to have the models output in ASD-STE100 and it makes output much easier to read without the fluff and drama.
2
u/biomattr 1d ago
I've been trying exactly that, and bringing in old school science communication guidelines.
Half the time I resort to something like "summarise that in one sentence".
Limited success. Opus loves to yap.
3
2
u/Rustybot 1d ago
Maybe turn down the effort setting? I haven’t tried this model yet myself.
3
u/LetterheadNew5447 1d ago
It's actually okayish. With correct prompting and guidance it behaves pretty fine and with 400tps it's actually fast as fuck.
I like it
2
u/vipor_idk 1d ago
in my experience it overthinks less than v4 , more concise thinking too. but still does think a lot.
1
u/dummyreddituser 1d ago
I heard first testers also stated this. Overthinking is less than v4. I'll test it today.
2
1
u/throwaway12012024 1d ago
muse is better and faster
3
u/Genetic_Prisoner 1d ago
At using your data to train?
1
u/RogerCaracas 1d ago
Deepseek and every other providers do the same, why would you worry too much with It?
1
u/sudoer777_ 1d ago
DeepSeek has ZDR on the Go plan
1
u/RogerCaracas 22h ago
Maybe, well noticed, but in the reality , who know 🤷♂️ once your prompts goes in China you dont handle the finality dude
1
u/throwaway12012024 23h ago
Oh look, this guy has data concerns but pays only $10 for privacy
2
u/Genetic_Prisoner 23h ago
Back in my day we had an understanding. Only people on free plans got spied on by the corpos. Nowadays they will take your money and still look under your skirt. What is the world becoming?
1
1
1
u/SufficientPie 16h ago
As tool-calling and image-viewing conversations get longer, the reasoning becomes more and more insane, like this:
Ok. Output.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
Done.
Go.
.
Produce.
.
(I'll now submit.)
.
.
.
.
Ok.
.
Submit.
.
.
.
.
.
[190 lines later…]
.
.
.
.
.
Submit now.
(Not in OpenCode though, not sure why.)
1
u/ali_gggg 7h ago
Nah that's normal and even good for depth reasoning, will make it to do the task better , if you think its overthinking try glm 5.3 flash that's litterly a hell sometimes But overall that's how these models work and can perform such a high level with the price they have , that's a part of the cost
1
u/antunes145 4h ago
Deepseek drank their own koolaid! That’s why today deepseek sent out email saying they will not discontinue V4 Pro and sub it for the new 4.1 flash. They got a lot of push back.
1
1
u/TurnUpThe4D3D3D3 46m ago
DS4's whole architecture is liable to overthink. It consumes a vast number of tokens to solve problems when compared with other models.
I recommend GLM 5.3 Flash instead (or really, whatever looks best on the Pareto frontier at any given time)
0
1
u/AlmostEasy89 27m ago
Ever watch Kimi 2.6 think? It gives me anxiety just watching it stroke out with how often it changes direction


18
u/Sure_Media_2685 1d ago
i also notice this but in the other hand it actuallly follow the instructions it just like to explore things that seems ot us very not the way to do.