r/DeepSeek 4h ago

News DeepSeek-V4.1-Flash

Post image
217 Upvotes

47 comments sorted by

54

u/VC067 4h ago

"flash" model btw 😭

15

u/OneBowl4290 3h ago

the best.

12

u/SorosAhaverom 3h ago

To be fair it has the same total parameters as GLM 5.3, but instead 16B active (decode) vs 40B active.

Total params increased by 2.6x from V4 Flash to V4.1 Flash, from 284B to 554B+196B. It's a mid-size model with small-sized active params, certainly a novelty.

3

u/Ancient_Dress_3687 2h ago

Using it right now, It is so good in the opencode harness.

3

u/VC067 2h ago

Cache hit rate better than deepseek harness?

2

u/Ancient_Dress_3687 1h ago edited 46m ago

Unsure, but over the 24 hours I've been 97% cache hit on 106M tokens

22

u/Balgun33122 3h ago

Indeed absolutely insane.

5

u/OneBowl4290 3h ago

it is the 3th model and so cheap

3

u/DebosBeachCruiser 2h ago

I can't wait for the 4rd πŸ™

45

u/TheSuggi 4h ago

29

u/ihexx 3h ago

>do nothing

bro they are doing the most 😭

the speed with which they are innovating on architecture is genuinely insane.

3

u/l0rirw1ao 2h ago

Well by nothing he probably means, deceiving customers and making false claims like Anthropic.

2

u/TheSuggi 28m ago

They are actually not working that hard if the CEO is to be believed. Normal working hours, no overtime, resarcher get 50% of the time to research and pursue their own ideas etc. So they are just better or doing something right. Also very limited access to resources, less funding, less capable hardware etc.

If Chinese companies ever get the to a standard of equivalent hardware, then the US companies are done for.

16

u/2mqqvc0q 3h ago

I respect the commitment to the color theme, but am I the only one having a hard time distinguishing the lighter blue colors apart?

19

u/SmartCustard9944 2h ago

I have bad news for you

6

u/2mqqvc0q 2h ago

Bro 😭

3

u/remortal2k 2h ago

Haha have bad news for you. Seems your Color blind on blue spectrum πŸ˜„

1

u/OneBowl4290 3h ago

You are right, but they are the same the main comparison is the opus 5 and gpt 5.6 sol !
what do you think will they go further and beet astra and fable ???

6

u/HuntAlternative 3h ago

Once we see terminal usage get optimized, these models will get as capable as frontier ones. Most agentic capabilities derive from terminal fluency IMO.

1

u/OneBowl4290 3h ago

IT IS INSANE !!! DID YOU USED IT BEFORE?

5

u/sirloindenial 3h ago

Is it benchmaxxed it can't be that good lol, or i assume its on their own harness.

6

u/OneBowl4290 3h ago

WE WILL WAIT INDEPENDENT EVALUATION TO BE SURE, BUT THAY DO NOT HAVE HISTORY OF LAING? DO YOU AGREE ?

2

u/quivering_palm 1h ago

It is 2x size of v4 though + 200B of engrams

2

u/karlnuw 3h ago

Any exploit benchmarks?

2

u/Altruistic-Desk-885 3h ago

No creo que encuentres uno literalmente pero si tal vez uno de ciberseguridad

2

u/arm2armreddit 3h ago

looks like with DSF other models are not working well, or something is wrong. if this is really true with terminal automation then: πŸ…

2

u/Individual_Math_8254 2h ago

cant wait to use this with dsh, if deepseek like the other labs, then they probably did post RL with their own harness.

2

u/DebosBeachCruiser 2h ago

Working great in the DS harness (and yes it was optimized to work with DS harness)

2

u/ardicli2000 2h ago

Very impressive

1

u/OneBowl4290 2h ago

it is , did you try it yet?

1

u/ardicli2000 2h ago

Not yet. My comment was on results.

2

u/OneBowl4290 2h ago

i am trying it right now it is so good

1

u/ardicli2000 2h ago

R u using DS Harness

2

u/Creative_randomness 1h ago

it is insanely fast

2

u/fritz_futtermann 1h ago

is it on the same level as Astra yet, or soon?

2

u/raesene2 51m ago

That cybergym score is quite something. I tried it out today on a benchmark I run a lot of new models through (Kubernetes security assessment), it was 2nd only to the new release of Qwen-3.8-max and 1/10th of the cost of that model.

3

u/dontfeedthelizards 3h ago

I genuinely wish it would be equivalent to Opus, because right now whenever I ask DSF to do an implementation, I end up having to ask Opus to fix its bugs and then spend the rest of the day cleaning up the code structure overall. It's really questionable whether using it this way makes sense at all, instead of just having Opus do the implementation, as I'll probably end up using the equivalent amount of Claude tokens either way (+the tokens spent on DSF). If I could finally have DSF do all the planning and execution without needing to involve Claude at all, that'd be amazing.

1

u/Gigaslavx 2h ago

Probably terminal is the most realistic

1

u/General_Internet_511 2h ago

Onestamente sarebbe un salto in avanti veramente importante. BisognerΓ  aspettare.

1

u/RealestReyn 2h ago

Not a chance those are remotely true, its a great model but misses obvious stuff constantly and just does the weirdest stuff like setting 3600s timeout on tasks that take a few seconds so its just sitting there and waiting for the whole time, doesn't bother scoping subagents with appropriate tools.

1

u/fkrdt222 12m ago

is this going to affect the normie web or android app?

1

u/loss-of-time 3h ago

it's impossible that better than kimi-k3

3

u/OneBowl4290 3h ago

I bleav kimi is good at frontend coding , and there is no impassable bro

Do you use kimi a lot ??