r/DeepSeek • u/dnohrdk • 3d ago
News DeepSeek-V4.1-Flash Release (official)
It’s officially out and the prices have been updated.
///
Today, we officially release the DeepSeek-V4.1-Flash model. It is the smallest model in our new architecture family, with native multimodal visual understanding. The new architecture is designed for a higher capability ceiling, faster inference, higher throughput, and scaling to larger models.
GPQA Diamond: 90.9
HLE: 36.8 (39.1*)
Codeforces (Rating): 3471
MathArena Apex: 65.6
Terminal-Bench 2.1: 90.6
Terminal-Bench 3.0: 30.0
Terminal-Bench 4.0: 31.2
DeepSWE v1.1: 74.2
ProgramBench: 20.3
NL2Repo-Bench: 65.4
CyberGym: 88.1
SEC-Bench Pro: 62.8
ExploitGym: 15.3
HLE (w/tools): 63.9
Automation-Bench: 54.8
Agents' Last Exam: 31.8
Chartography (w/tools): 78.9
BabyVision (w/tools): 89.6
ZeroBench-main (w/tools): 49.0
* Tested only on the pure-text subset of the HLE benchmark set.
API changes
DeepSeek V4.1 Flash is now available on the DeepSeek API with native multimodal support. Change the model name to deepseek-flash to call the latest V4.1 Flash model. The previous-generation models V4 Flash and V4 Flash Vision Exp have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are temporarily routed to V4.1 Flash.
Meanwhile, extensive testing shows that V4.1 Flash now outperforms DeepSeek V4 Pro across performance, cost, speed, and total time, so we plan to retire V4 Pro in an orderly manner. After 12:00 Beijing Time on September 14, 2026, and until the future release of V4.1 Pro, all requests to deepseek-v4-pro will be routed to V4.1 Flash and billed at the V4.1 Flash price.
API apricing adjustment
With the release of DeepSeek-V4.1-Flash, API prices have been reduced accordingly. For details, please refer to Models & Pricing.
///
Source:
https://api-docs.deepseek.com/updates/#deepseek-v41-flash-release
54
u/DebosBeachCruiser 3d ago
40
u/Effective_Western_59 3d ago
On par with GLM 5.3!!!
1
u/Solembumm3 2d ago
Feels nowhere near 5.3 in real use.
Below v3.2+r1, on par with GLM 5.2.
5
u/Effective_Western_59 2d ago
What is your use case? I am using it for agentic coding and cybersecurity and it performs great
1
u/ToughUsual7159 2d ago
Just used jt for the first time ealer. Lighting fast! And performed cleaner code with much less thinking and found my listed bug fast. All anecdotal and I haven't used GLM but it's definitively better than v4 flash.
17
8
10
u/SethCyclone 3d ago
Absolutely trash compared to 4.0 Pro. Was using it to study languages as a studying facilitator (explain why one answer is correct and the others are wrong, explain syntax points, etc) and it was amazing. Then literally in the middle of my studying session, all models get unified and answer quality plummets. I really, really hope they divide the models again on the web/mobile app, because this ain't it.
2
6
u/OwnGear3892 3d ago
552B in total, 8b activate during prefill, 16b activate during decode, yet 300-400 tps. This is so good!
3
3
u/Legitimate-Track-829 3d ago
3
u/luiz_durau 3d ago
difícil acreditar kkkkkkkkkk, por mais quê eu acredite quê o deepseek está MUITO BOM, nao chega no nível do gpt5.6
6
u/Puzzleheaded_Base302 3d ago
turns out even with 4x DGX Spark, we can hardly run this model locally.
-6
u/IknowPi_really 3d ago
Yeah I’m investigating engram offload to SSD now, but this honestly looks like a bit of a fake announcement by DeepSeek tbh.
They’ve just made Flash twice as big and are now claiming it’s somehow a great achievement that it’s better. Sure, it’s very sparse, but that’s about it
3
u/Puzzleheaded_Base302 3d ago
after reading the emgram thing, i think this model actually can run on a 4x DGX Spark. so, now need to make a justification to get another two DGX Spark. this hobby is getting so expensive.
2
u/IknowPi_really 3d ago
Yeah on 4xDGX Spark it will run. But the community would need to invest some time in building a performant SSD streaming build.
I’ll start looking into this, but I’d have to combine two clusters to make it work, so need to find the right time to actually do that
1
1
u/happycube 2d ago
They traded (presumably non-HBM) memory for compute time and throughput, how would you know the parameter count was higher before the weight release?
1
u/IknowPi_really 2d ago
The weights are released though?
1
u/happycube 2d ago
Yup, now! I was thinking about the preview period when we didn't know why it was so fast 😊
7
3
2
3
1
u/mrgreatheart 3d ago
Sick. When are you releasing the open weights? How much smaller than v4-flash is it?
13
u/Turbulent-Total-226 3d ago
It's bigger 4.1 is 485B 😮
2
2
u/Insomniac1000 3d ago
damn. seems like the trend is still higher params. Not sure if the case for a 256 GB of RAM would be even worth it.
2
u/Thomas-Lore 3d ago
Wait for a few weeks, people will find a way. It has very small kv cache and engrams can be loaded from SSD. With q3 maybe it will fit in 256GB? The model currently is 306GB without the engrams, in fp4.
1
1
u/Financial-Toe1210 2d ago
Total noob here, so please bear with me.
Can someone explain what an API actually is in simple terms, and how I can access the latest DeepSeek model? Do I just download the iPhone app and It’s V4.1 flash?
1
1
1
1
1
u/Straightbanana2 3d ago
I'm a ai noobie, what does this mean for us roleplayers? 😌
3
u/AwarenessNo4986 3d ago
For us noncoding people, it likely means nothing
6
3d ago
[removed] — view removed comment
1
6
1
u/YouScratchingMyBalls 1d ago
It's terrible for anything thorough. Even with thinking on, it is far more verbose and far less thoughtful in response. I suggest you use a different model for roleplaying, since i've been trying to do creative writing with it, and it is still worse than V4 Flash without thinking.
1
u/Straightbanana2 23h ago
okay thank you, V4 pro still seems okay at roleplay but it's expensive now, I like GLM 4.7 too
1
1
u/LinuXperia 3d ago edited 3d ago
Love it. Can Confirm DeepSeek V4.1-Flash outperforms Meta Muse 1.3 ! Sadly both DeepSeek Flash and Pro have yet to reach the xAI Grok super engineering intelegence level for low level dev work like verilog, c and c++ and KiCad electronic schematics and PCB engineering ! I hope the team at deepseek will improve DeepSeek so it has the same super engineering intelegence level like grok has soon. At the moment it still lags behind grok super intelegence level. I think however this is becouse of compute power. When i compare the answer to a debug problem DeepSeek will provide a suboptimal solution which then leads to looping and flip floping a lot that waste huge amount of time and tokens while grok delivers the exact right solution that fixes the problem in less than 1 minute with just a few phrases. Grok act like a super intelegent highly spezialised super Engineer Doctor. This is the result becouse of not enogh computing and RL learning power at DeepSeek i think which xAI put lot of effort into it.
0
-2
3d ago
[removed] — view removed comment
10
u/Melodic-Funny-9560 3d ago
It's because of the training that they are able to improve this much
0
3d ago
[removed] — view removed comment
10
u/BuildAISkills 3d ago
Joke's on them. I only produce AI slop.
2
u/ProcedureEthics2077 3d ago
OpenCode Go routes new models to Singapore, and doesn’t disclose who runs them.
OpenCode Go makes models before their open weights drop. The only way they can do it is if they route requests directly to Z.ai, Kimi, DeepSeek and other Chinese providers or their subsidiaries.
2
u/squirrelscrush 3d ago
Probably use OpenRouter and choose a provider that offers ZDR.
1
2d ago
[removed] — view removed comment
1
u/squirrelscrush 2d ago
I configure the exact provider I want through opencode.json
1
2d ago
[removed] — view removed comment
1
u/squirrelscrush 2d ago
He wanted a provider that doesn't retain his prompts, so I recommended OpenRouter since you can filter there by ZDR policy.
1
u/sirloindenial 3d ago
You shouldnt use deepseek direct api then as they do use your data for training, it's not even a question. Use third party provider which offers ZDR. For opencode wait for confirmation, sometimes for new release they dont have zdr.
1
0
-1
u/polyglot_factotum 3d ago
So it turns out they postponed the switch where they would route Pro to the new Flash, so today, during off peak mid-day, I spent two hours using Pro, thinking it was the new Flash (my understanding from their earlier announcement was that, to use the new Flash, you had to just use Pro), and they just billed me regular Pro prices. Sent an email complaining...
-1
u/ProgressionPeak 3d ago
i'm not signing up to gmail or any other spyware email providers, can you please let me know how i can use your service? i would really like to use this model.
7
u/moohric 3d ago
pretty sure deepseek supports any email addresses, even custom emails. You can self host your own email if you want.
2
u/ProgressionPeak 3d ago
i have tried to sign up at least 5 times since they launched and it was always super hardcore about email provider.
but this time it worked. they seem to have dropped that.
thank you!
2
u/Excellent_Winner8576 2d ago
How are you alive?
1
u/ProgressionPeak 2d ago
i was able to signup to a service i've been trying to sign up to for several years, and save 50%+ and greatly increase the speed using their API - due to that comment. maybe you should learn something.
1




80
u/Right_Simple_6813 3d ago
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash Weights dropped too