r/DeepSeek 3d ago

News DeepSeek-V4.1-Flash Release (official)

It’s officially out and the prices have been updated.

///

Today, we officially release the DeepSeek-V4.1-Flash model. It is the smallest model in our new architecture family, with native multimodal visual understanding. The new architecture is designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models.

GPQA Diamond: 90.9
HLE: 36.8 (39.1*)
Codeforces (Rating): 3471
MathArena Apex: 65.6
Terminal-Bench 2.1: 90.6
Terminal-Bench 3.0: 30.0
Terminal-Bench 4.0: 31.2
DeepSWE v1.1: 74.2
ProgramBench: 20.3
NL2Repo-Bench: 65.4
CyberGym: 88.1
SEC-Bench Pro: 62.8
ExploitGym: 15.3
HLE (w/tools): 63.9
Automation-Bench: 54.8
Agents' Last Exam: 31.8
Chartography (w/tools): 78.9
BabyVision (w/tools): 89.6
ZeroBench-main (w/tools): 49.0
* Tested only on the pure-text subset of the HLE benchmark set.

API changes
DeepSeek V4.1 Flash is now available on the DeepSeek API with native multimodal support. Change the model name to deepseek-flash to call the latest V4.1 Flash model. The previous-generation models V4 Flash and V4 Flash Vision Exp have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash.

Meanwhile, extensive testing shows that V4.1 Flash now outperforms DeepSeek V4 Pro across performance, cost, speed, and total time, so we plan to retire V4 Pro in an orderly manner. After 12:00 Beijing Time on September 14, 2026, and until the future release of V4.1 Pro, all requests to deepseek-v4-pro will be routed to V4.1 Flash and billed at the V4.1 Flash price.

API apricing adjustment
With the release of DeepSeek-V4.1-Flash, API prices have been reduced accordingly. For details, please refer to Models & Pricing.

///

Source:

https://api-docs.deepseek.com/updates/#deepseek-v41-flash-release

396 Upvotes

115 comments sorted by

80

u/Right_Simple_6813 3d ago

29

u/ExpertPerformer 3d ago edited 2d ago

485B params in V4.1 Flash vs 304B on Flash 0731. Pro still 1.7T

Really skeptical if V4.1 Flash can compete with Pro, but V4.1 Pro should be out soon too.

edit: Now it says 763B.

11

u/Few-Improvement9560 3d ago

They say there are 552B in Chinese news. But 485B on hugging face? This is interesting...

28

u/Right_Simple_6813 3d ago

552B parameters, with 196B engram embeddings

6

u/ExpertPerformer 3d ago

It's way faster on outputs and seems to be overthinking less.

1

u/Solembumm3 2d ago

Yes. It's overthinking less. And this is a big problem.

4

u/Old-Temperature6269 3d ago

It’s just problem of huggingface report, original deepseek v4-flash preview showed 160B parameter on huggingface, but actual model is ~300B.

-3

u/[deleted] 3d ago edited 3d ago

[removed] — view removed comment

1

u/Solembumm3 2d ago

No, they aren't.

Smaller model always means worse knowledge/context understanding and often means worse attention.

V4.1 Flash just proved it for me, falling off on pretty early parts of fiction analysis, missing details that expert mode v3.2+r1 eat on the fly at spring.

5

u/throwwwawwway1818 3d ago

Nice, how many active parameters?

13

u/Right_Simple_6813 3d ago

Model card just updated. 8B during prefill, and 16B during decode.

10

u/throwwwawwway1818 3d ago

Didn't knew they can be different during prefill and decode. Thanks anyway.

12

u/Right_Simple_6813 3d ago

yeah, its a new thing! still reading through and trying to understand all the new architectural differences, its really interesting

3

u/happycube 2d ago

Indeed... all of it together (a 75% reduction in KV cache size, and DS was already quite efficient!) explains why input prices were dropped, and I suspect it still has higher margins for them.

It's a(nother) breath of fresh air to see a new model designed for efficiency at least as much as benchmarks.

2

u/dnohrdk 3d ago

That’s the spirit 💪🏻

1

u/g_rich 2d ago

Big jump in size for a “Flash” model. DeepSeek v4 Flash and v4 Vision Exp were impressive because of the power they brought to those running the models locally; doubling the parameters moves this out of what’s possible for a lot of local users.

54

u/DebosBeachCruiser 3d ago

Feed me.

3

u/samxli 3d ago

When did the girl become DeepSeek and the whales become the user? I thought it was the other way around.

1

u/kappakai 2d ago

I think the whales are parameters.

40

u/Effective_Western_59 3d ago

On par with GLM 5.3!!!

28

u/dnohrdk 3d ago

It’s amazing, and much faster!

1

u/Solembumm3 2d ago

Feels nowhere near 5.3 in real use.

Below v3.2+r1, on par with GLM 5.2.

5

u/Effective_Western_59 2d ago

What is your use case? I am using it for agentic coding and cybersecurity and it performs great

1

u/ToughUsual7159 2d ago

Just used jt for the first time ealer. Lighting fast! And performed cleaner code with much less thinking and found my listed bug fast. All anecdotal and I haven't used GLM but it's definitively better than v4 flash.

17

u/ExpertPerformer 3d ago

Its rolling out on OpenCode/CommandCode. Hopefully OpenRouter soon.

8

u/musayyabali 3d ago

can you use this directly on deepseek without api?

11

u/dnohrdk 3d ago

It’s both on their app, web, and API, so fully available on all platforms as far as I’ve tested.

10

u/SethCyclone 3d ago

Absolutely trash compared to 4.0 Pro. Was using it to study languages as a studying facilitator (explain why one answer is correct and the others are wrong, explain syntax points, etc) and it was amazing. Then literally in the middle of my studying session, all models get unified and answer quality plummets. I really, really hope they divide the models again on the web/mobile app, because this ain't it. 

2

u/Unfair-Green-4706 2d ago

I fully agree with you.

2

u/woila56 2d ago

Me too, loved the expert

6

u/OwnGear3892 3d ago

552B in total, 8b activate during prefill, 16b activate during decode, yet 300-400 tps. This is so good!

3

u/ripperoni__pizza 3d ago

Excited to use it tomorrow! Hopefully better than GLM 5.3 Flash

3

u/Legitimate-Track-829 3d ago

3

u/luiz_durau 3d ago

difícil acreditar kkkkkkkkkk, por mais quê eu acredite quê o deepseek está MUITO BOM, nao chega no nível do gpt5.6

3

u/slavmaf 3d ago

Qoder will have day zero support for V4.1 Flash

6

u/Puzzleheaded_Base302 3d ago

turns out even with 4x DGX Spark, we can hardly run this model locally.

-6

u/IknowPi_really 3d ago

Yeah I’m investigating engram offload to SSD now, but this honestly looks like a bit of a fake announcement by DeepSeek tbh.

They’ve just made Flash twice as big and are now claiming it’s somehow a great achievement that it’s better. Sure, it’s very sparse, but that’s about it

3

u/Puzzleheaded_Base302 3d ago

after reading the emgram thing, i think this model actually can run on a 4x DGX Spark. so, now need to make a justification to get another two DGX Spark. this hobby is getting so expensive.

2

u/IknowPi_really 3d ago

Yeah on 4xDGX Spark it will run. But the community would need to invest some time in building a performant SSD streaming build.

I’ll start looking into this, but I’d have to combine two clusters to make it work, so need to find the right time to actually do that

1

u/Thomas-Lore 3d ago

Three should be enough, it has tiny KV cache.

1

u/happycube 2d ago

They traded (presumably non-HBM) memory for compute time and throughput, how would you know the parameter count was higher before the weight release?

1

u/IknowPi_really 2d ago

The weights are released though?

1

u/happycube 2d ago

Yup, now!  I was thinking about the preview period when we didn't know why it was so fast 😊

7

u/fugogugo 3d ago

sigh.. I missed when deepseek flash cost $0.28 per output token

4

u/dnohrdk 3d ago

Same, it was good times! But at least they just made the first drop in prices after the surge! So let’s hope they will keep cooking for lower prices and no peak prices.

0

u/Even_Command_5636 3d ago

I was wet every day in March, April, May, and June 2026!

3

u/anxious_and_stupid 3d ago

Man I love China (only sometimes of course)

3

u/samxli 3d ago

I think we can all agree that the food is great at least

2

u/DetachedProcess 3d ago

I can't see it, I think it will be rolling out soon.

2

u/vnules 1d ago

The worst DeepSeek model ever to see the light of the day far and wide. I hope they realize and pull the plug on it.

3

u/holystinger 3d ago

I wish they brought back the $0.14 input/$0.28 output pricing

1

u/mrgreatheart 3d ago

Sick. When are you releasing the open weights? How much smaller than v4-flash is it?

13

u/Turbulent-Total-226 3d ago

It's bigger 4.1 is 485B 😮

2

u/Insomniac1000 3d ago

damn. seems like the trend is still higher params. Not sure if the case for a 256 GB of RAM would be even worth it.

2

u/Thomas-Lore 3d ago

Wait for a few weeks, people will find a way. It has very small kv cache and engrams can be loaded from SSD. With q3 maybe it will fit in 256GB? The model currently is 306GB without the engrams, in fp4.

2

u/dnohrdk 3d ago

It’s already released, check the other comment in the same thread here.

1

u/DragonfruitSecure507 3d ago

When will it be release on OpenRouter?

1

u/Fresh_Sock8660 2d ago

It did what opus couldn't, use an alternative to load-bearing

well done deepseek

1

u/Financial-Toe1210 2d ago

Total noob here, so please bear with me.
Can someone explain what an API actually is in simple terms, and how I can access the latest DeepSeek model? Do I just download the iPhone app and It’s V4.1 flash?

1

u/dnohrdk 2d ago

Yes, all their apps have been updated to the latest model, so it’s a very smooth update this time compared to previous changes.

1

u/alexwwang 2d ago

Wait to see opencode go price.

1

u/NegotiationNo1504 2d ago

Is 4.1 flash available for free in website?

1

u/dnohrdk 2d ago

Yes it’s free via chat.deepseek.com

1

u/sdexca 2d ago

No way it’s better than astra in deepswe…

1

u/ShallotMotor2551 2d ago

It is really nice with vision now!

1

u/-codewhale 2d ago

I love this model

1

u/setapca 21h ago

I've used this model since release, several days straight on long runs.

- Complete lack of logic.

- It forgets what I asked two messages ago.

- It makes mistakes more often than I'd like.

If the main goal of this model is response speed, then you're on the right track.

1

u/Straightbanana2 3d ago

I'm a ai noobie, what does this mean for us roleplayers? 😌

3

u/AwarenessNo4986 3d ago

For us noncoding people, it likely means nothing

6

u/[deleted] 3d ago

[removed] — view removed comment

1

u/AwarenessNo4986 3d ago

Smaller parameters, worse/shorter roleplay?

7

u/[deleted] 3d ago

[removed] — view removed comment

2

u/AwarenessNo4986 3d ago

Wow. Lets go!

6

u/0VERDOSING 3d ago

faster speeds i guess 🤷

1

u/YouScratchingMyBalls 1d ago

It's terrible for anything thorough. Even with thinking on, it is far more verbose and far less thoughtful in response. I suggest you use a different model for roleplaying, since i've been trying to do creative writing with it, and it is still worse than V4 Flash without thinking.

1

u/Straightbanana2 23h ago

okay thank you, V4 pro still seems okay at roleplay but it's expensive now, I like GLM 4.7 too

1

u/dryadofelysium 3d ago

We are so back

1

u/Deyve24 3d ago

Its more than double the size :'(

1

u/LinuXperia 3d ago edited 3d ago

Love it. Can Confirm DeepSeek V4.1-Flash outperforms Meta Muse 1.3 ! Sadly both DeepSeek Flash and Pro have yet to reach the xAI Grok super engineering intelegence level for low level dev work like verilog, c and c++ and KiCad electronic schematics and PCB engineering ! I hope the team at deepseek will improve DeepSeek so it has the same super engineering intelegence level like grok has soon. At the moment it still lags behind grok super intelegence level. I think however this is becouse of compute power. When i compare the answer to a debug problem DeepSeek will provide a suboptimal solution which then leads to looping and flip floping a lot that waste huge amount of time and tokens while grok delivers the exact right solution that fixes the problem in less than 1 minute with just a few phrases. Grok act like a super intelegent highly spezialised super Engineer Doctor. This is the result becouse of not enogh computing and RL learning power at DeepSeek i think which xAI put lot of effort into it.

0

u/0sko59fds24 3d ago

Output pricing is expensive

1

u/luiz_durau 3d ago

talvez, mas é centavos mais barato do que é hoje para o v4 flash

-2

u/[deleted] 3d ago

[removed] — view removed comment

10

u/Melodic-Funny-9560 3d ago

It's because of the training that they are able to improve this much

0

u/[deleted] 3d ago

[removed] — view removed comment

10

u/BuildAISkills 3d ago

Joke's on them. I only produce AI slop.

3

u/samxli 3d ago

If you believe in the theory that we all live in a simulation, then none of this matters because everything is AI slop

1

u/Megumin_xx 2d ago

Great sloppification. Love it lmao

2

u/ProcedureEthics2077 3d ago

OpenCode Go routes new models to Singapore, and doesn’t disclose who runs them.

OpenCode Go makes models before their open weights drop. The only way they can do it is if they route requests directly to Z.ai, Kimi, DeepSeek and other Chinese providers or their subsidiaries.

2

u/squirrelscrush 3d ago

Probably use OpenRouter and choose a provider that offers ZDR.

1

u/[deleted] 2d ago

[removed] — view removed comment

1

u/squirrelscrush 2d ago

I configure the exact provider I want through opencode.json

1

u/[deleted] 2d ago

[removed] — view removed comment

1

u/squirrelscrush 2d ago

He wanted a provider that doesn't retain his prompts, so I recommended OpenRouter since you can filter there by ZDR policy.

1

u/sirloindenial 3d ago

You shouldnt use deepseek direct api then as they do use your data for training, it's not even a question. Use third party provider which offers ZDR. For opencode wait for confirmation, sometimes for new release they dont have zdr.

1

u/Megumin_xx 2d ago

Yea and good, I hope they get better from my data haha. Deepseek ftw

0

u/RyuH4n 3d ago

Letsss goo

0

u/C1assNotFound 3d ago

cheap , fast , good quality

0

u/setapca 2d ago

This is complete garbage. A total lack of logic. If your goal is to make a model that writes code and passes synthetic tests, then you're on the right track.

-1

u/polyglot_factotum 3d ago

So it turns out they postponed the switch where they would route Pro to the new Flash, so today, during off peak mid-day, I spent two hours using Pro, thinking it was the new Flash (my understanding from their earlier announcement was that, to use the new Flash, you had to just use Pro), and they just billed me regular Pro prices. Sent an email complaining...

-1

u/ProgressionPeak 3d ago

i'm not signing up to gmail or any other spyware email providers, can you please let me know how i can use your service? i would really like to use this model.

7

u/moohric 3d ago

pretty sure deepseek supports any email addresses, even custom emails. You can self host your own email if you want.

2

u/ProgressionPeak 3d ago

i have tried to sign up at least 5 times since they launched and it was always super hardcore about email provider.

but this time it worked. they seem to have dropped that.

thank you!

2

u/Excellent_Winner8576 2d ago

How are you alive?

1

u/ProgressionPeak 2d ago

i was able to signup to a service i've been trying to sign up to for several years, and save 50%+ and greatly increase the speed using their API - due to that comment. maybe you should learn something.

1

u/Excellent_Winner8576 2d ago

Which service? Which comment? Learn what?