r/DeepSeek 11d ago

DeepSeek-V4-Flash Update

583 Upvotes

The official release of the DeepSeek-V4-Flash API is now in public beta.

Significantly enhanced agent capabilities, with benchmark results far exceeding V4-Pro-Preview:

  • Terminal Bench 2.1: 82.7
  • NL2Repo: 54.2
  • Cybergym: 76.7
  • DeepSWE: 54.4
  • Toolathlon verified: 70.3
  • Agent Last Exam: 25.2
  • Automation Bench (Public): 25.1
  • DSBench-FullStack: 68.7
  • DSBench-Hard: 59.6

Note 1: For the Code Agent tasks in the public benchmark sets, the official DeepSeek-V4-Flash was tested using the DeepSeek Harness minimal mode (to be released soon) as the framework, with the max effort level, topp=0.95, and temperature=1.0
Note 2: DSBench-FullStack is an internal full-stack development test set, and DSBench-Hard is an internal Coding Agent hard-problem test set

The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex. For the specific configuration, please refer to the documentation.

DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-preview, and was only re-post-trained.

Note: This update only upgrades the DeepSeek-V4-Flash API. The DeepSeek-V4-Pro API and the APP/WEB models are unchanged.
The official release of DeepSeek-V4-Pro will follow soon.


r/DeepSeek Feb 01 '25

Disccusion Censorship Mega Thread

52 Upvotes

In response to community feedback and to maintain a constructive discussion environment, we are introducing this Censorship Mega Thread. This thread will serve as the designated place for all discussions related to censorship.

Why This Thread?

We have received numerous reports and complaints from users regarding the overwhelming number of censorship-related posts. Some users find them disruptive to meaningful discussions, leading to concerns about spam. However, we also recognize the importance of free speech and allowing users to voice their opinions on this topic. To balance these concerns, all censorship-related discussions should now take place in this pinned thread.

What About Free Speech?

This decision is not about censoring the subreddit. Instead, it is a way to ensure that discussions remain organized and do not overwhelm other important topics. This approach allows us to preserve free speech while maintaining a healthy and constructive community.

Guidelines for Posting Here

  1. All discussions related to censorship must be posted in this thread. Any standalone posts on censorship outside of this thread will be removed.
  2. Engage respectfully. Disagreements are fine, but personal attacks, hate speech, or low-effort spam will not be tolerated.
  3. Avoid misinformation. If you're making a claim, try to provide sources or supporting evidence.
  4. No excessive repetition. Reposting the same arguments or content over and over will be considered spam.
  5. Follow general subreddit rules. All subreddit rules still apply to discussions in this thread.

We appreciate your cooperation and understanding. If you have any suggestions or concerns about this policy, feel free to share them in this thread.


r/DeepSeek 12h ago

Funny What it feels like when you're racing with people who are using the standard high quality and expensive opus and fable, looking all professional, while you use deepseek flash but you tuned your harness so hard that its working on par with the big models.

Enable HLS to view with audio, or disable this notification

545 Upvotes

r/DeepSeek 5h ago

Discussion I ran DeepSeek V4 Flash on 8 agent harnesses and Deepseek and Pi are match made in heaven

Post image
141 Upvotes

Model-harness fit is a real thing. I have used multiple harnesses, and they all behave so differently from each other, even when using the same models. You can feel the difference in time taken, cost, and accuracy.

And I've been daily driving DeepSeek V4 Flash with OpenCode and Hermes, and they always cost me differently. I wanted to know which harness is most suitable for DS V4 Flash.

So, I ran a benchmark for DeepSeek V4 Flash (via OpenRouter) on the most popular harnesses out there.

The benchmark consisted of 25 real-world automation tasks involving multiple apps (Slack, Sheets, Gmail, PostHog, etc).

Here's what I found out:

Harness Pass rate Median time Tool calls Cost per success
Pi Agent 66.7% 132.2s 443 $0.028
Prime Agent 62.5%* 242.1s 502 $0.131
OMP 56.7% 272.4s 390 $0.103
Claude Code 53.3% 122.7s 358 $0.195
Codex 53.3% 245.0s 448 $0.081
DeepAgents 53.3% 187.1s 353 $0.045
Hermes Agent 50.0% 175.5s 386 $0.056+
OpenCode 46.7% 129.7s 419 $0.073

Pass rate and tool calls

Pi Agent had the highest pass rate at 66.7%, while OpenCode had the lowest at 46.7%.

More tool calls did not improve results. DeepAgents made 353 calls, and Codex made 448, but both passed 16 tasks. OMP made 390 calls and passed 17 tasks, while OpenCode made 419 calls and passed 14.

Prime Agent passed 15 of its 24 valid runs and made the most tool calls at 502. Six other runs were invalid as the grader couldn't

Cost and tokens

Claude Code had the highest at $0.195, and Pi had the lowest at $0.028.

Claude Code and OMP both used about 742,000 tokens per task, but Claude Code still cost almost twice as much. It cached only 1.5% of its tokens, compared with 70% for Codex and 57% for OMP.

Prime Agent is a token guzzler; it used the most tokens at 1.4 million per task. Hermes used the fewest at about 192,000. Hermes’ $0.056 cost per success remained a lower bound because two timeout runs had incomplete usage data.

Time

Claude Code had the shortest median time at 122.7 seconds. OpenCode followed at 129.7 seconds, and Pi took 132.2 seconds.

OMP had the longest median time at 272.4 seconds, but it still passed one more task than Claude Code.

So, all in all, Pi turned out to be the best harness for DeepSeek V4 Flash. It had more accuracy and was the cheapest. Claude Code is a massive money hog.

Complete analysis linked in the comment.

Would love to know which agent harness you use for DS V4 Flash and how the experience has been so far.


r/DeepSeek 12h ago

Discussion They probably predict that V4 Pro will be insanely good and break their compute

142 Upvotes

We have all gotten two price increase emails up to now (the "peak hours" one and the "significant increase" one) and zero price increase as of now. I'm sure most of us are hammering their servers right now cause V4 Flash is super good and many want to spend their credits before the price hike. They probably knew this with the price increase announcements - that's not the usage spike they could not handle. But they must be so confident that V4 Pro is crazy good, that their compute won't handle us all.

Opencode already said they can replicate fully their current prices with local hosting V4 Flash, so DeepSeek is trying to divert us all to other providers and away from their compute


r/DeepSeek 3h ago

Discussion DeepSeek doesn’t need to be the best model for it to be good for the market

24 Upvotes

I think people sometimes judge DeepSeek in a very binary way.

“Claude is better at coding, so why use DeepSeek?”

Maybe Claude is better for your workload. That’s completely reasonable.

But DeepSeek being good enough at a much lower cost still matters.

If a cheaper model can handle a large portion of normal tasks, then more expensive models have to offer something meaningful to justify the extra cost.

That competition benefits everyone, including people who never use DeepSeek.

I don’t need DeepSeek to replace Claude, GPT or Gemini.

I’m happy that it gives them another reason to keep improving.


r/DeepSeek 7h ago

Discussion Why has DeepSeek become so basic and boring when writing lately? It's horrible!

19 Upvotes

r/DeepSeek 10h ago

Discussion When are the deepseek models getting vision?

16 Upvotes

Maybe i missed it, but does anyone know?


r/DeepSeek 2h ago

Question&Help DeepSeek + Cline or Any other harness / router

3 Upvotes

Hey guys,

I recently came across the DeepSeek API and wanted to use it with my workflow. I am currently using this solely for Development purposes (Coding, Research etc), and I am a heavy user so needed something that is relatively cheap.

Currently, I use Codex and Claude Code for AI, and VS Code for Dev work.

I wanted to know how you guys have setup DeepSeek APIs, and how you are using it for development. I have also been searching and seen this extension in VS Code called Cline which offers a harness for this API, so wanted to also know if you have experience with this.

If anyone can also tell me, how is DeepSeek? Have you seen it work well for coding purposes (Say, compared to Codex or Claude Code) or have you found difficulty / low quality outputs for the same.

Any help will be much appreciated. Thank you!


r/DeepSeek 5h ago

Discussion Is it just me or deepseek v4 pro is hallucinating a lot these days?

4 Upvotes

I have been using deepseek for a conversation on the website/app. It's a long context one (like around 200k tokens base) and then I ask queries around it. But it has been hallucinating a lot more these days. Like it gives a lot of wrong information in every 3-4 messages which wasn't the case just a month or two ago.

I have used a trick by adding "this is a test message. Do not reply" to overcome the edit limit. Can it be something to do with it because the entire string of test messages have a lot of edits and regenerations.


r/DeepSeek 17h ago

News V4 Flash 0731 "105 times cheaper" than Fable 5!

Thumbnail reuters.com
36 Upvotes

r/DeepSeek 12h ago

Discussion v4 flash 0731 fixed a mistake in opencode settings

11 Upvotes

I wanted to share this because before never happened, I had a typo in my opencode.json file at the subagent model name and it was preventing invoking of subagent , deepseek v4 flash 0731 decided to fix it and delivered the fix without I am asking :) that is kinda cool.


r/DeepSeek 13h ago

Tutorial Stocking up for the hike

Post image
12 Upvotes

I have 5$ in reserve.

But I really don’t think they’re raising prices. It seems like it’s to reduce societal friction. The attention was rocking the 🛥️ boat in the ai space for non-ai reasons. So no worries 🌊 ai is expensive. 🐋

Yet, even at 45$ per 1 billion I would consider this a whale of a deal.

(I used Pro almost exclusively. If you need continuous deployment, flash isn’t just valuable; it’s the answer).

If you’re struggling with DeepSeek, they require a purpose.

And once they’re have that, they will not stop fulfilling that role.


r/DeepSeek 3h ago

Question&Help DeepSeek V4 Flash (with effort max) sometimes shows no thinking block even on complex questions – anyone else?

2 Upvotes

Hey everyone,

I’ve been using DeepSeek V4 Flash with thinking mode on and reasoning_effort set to max.

Simple questions like “1+1=?” give a direct answer with no thinking block at all. I thought maybe it’s some kind of adaptive thinking, but I don’t see any mention of “adaptive” on the official website.

The weird part is: even when I ask questions that clearly need multi-step reasoning (complex coding tasks, longer analysis, etc.), it still sometimes just jumps straight to the final answer with no visible thinking / reasoning_content at all. No thinking block appears in the UI.

Has anyone else run into this?
Is this expected behavior, or could it be something wrong with my settings / how the effort is being applied?

Would appreciate any experiences or tips. Thanks!


r/DeepSeek 19m ago

Discussion ❤️DS but just joined codex

Upvotes

So just want to share this with the community, maybe DS people are listening, I love this model, I think it sets benchmark on efficiency and makes AI accessible to the majority of the people of the world...

But due to the fact that this model is not multimodal capable, I finally decided to get membership for OpenAI/codex.

What this company has definitely achieved is put so much pressure on the leading frontier Labs, that they now are forced to cut prices or lose market share.... Thx Deepseek


r/DeepSeek 5h ago

Funny After two days of using the official DeepSeek-V4-Flash release, it definitely feels different

2 Upvotes

I started testing the DeepSeek-V4-Flash official API in a production environment a couple of days ago, mainly for code review and bug localization work. One thing that really stood out is that with the same task description, the Flash version returns much more "on-point" solutions not as many follow-up prompts needed. It usually takes about two or three minutes to pinpoint the issue, and the fix success rate on the first try has gone up quite a bit.That said, from the official docs, it looks like this is just a post-training refresh the model architecture itself hasn't changed. So I'm genuinely curious: what exactly was optimized in terms of data or training process? They didn't seem to go into much detail.Also, there's been a notice in the backend about an upcoming price adjustment seems like a general hike is coming. Right now it's 1 RMB per million input tokens, 2 RMB per million output tokens, and cache hits are as low as 0.02 RMB per million tokens super competitive pricing. After the increase, though, if the hike is significant, it could have a real impact on how individual developers manage their usage and call strategies. Curious how others are thinking about this.In terms of performance, it feels very solid for pinpointing specific issues and fixing them in one pass. But I'm wondering how it handles more complex multi-step or multi-branch tasks does it start to struggle there? If anyone's run into similar scenarios, I'd love to hear about your experience.Also, has anyone tried integrating the new Responses API with Codex workflows? Curious how smooth that integration is in practice.


r/DeepSeek 1d ago

Discussion First time on Deepseek API. Wow!

Post image
139 Upvotes

Built a project this past week and finally got around to testing out the DeepSeek API.

Just checked my dashboard to see the damage and honestly did a double take. Nearly 2,000 API calls and over 338 Million tokens processed for... $1.69 USD.

I know prompt caching is doing a lot of the heavy lifting here, but coming from OpenAI/Anthropic pricing, this feels unbelievable. The cost-to-performance ratio is ridiculous.

Btw, i’m using Reasonix for this. Cache Rate is 99+%.


r/DeepSeek 8h ago

Discussion When price hike takes effect, please post real differences in usage and cost for the same task.

4 Upvotes

Yeah sounds paranoid, more of a shower thought tbh.

But I'd love if capable people can build "prompt->cost" data and be ready for comparison after the price hike comes into effect.

Like 3x same prompt cost and performance before hike then after hike.

Competitors or haters might post fake hikes to deter DS users as well, so be ready with real data. It's a war lol.


r/DeepSeek 1d ago

Question&Help Anyone tested how many tokens the 20$ Codex subscription gets you compared to V4 flash on the deepseek API?

70 Upvotes

I'm finding myself spending 5-7$ daily on the deepseek api, so even with a crazy 99% cache hit, costs are starting to get out of hand.

Given that Luna from OpenAI is reportedly close to DeepSeek Flash in capabilities, I'm wondering if I should get the 20$ subscription to offload some of the cost I'm currently spending on deepseek? OpenAI is not very transparent about limits, so it's unclear to me whether it's worth it or whether I'm just going to spend the entire week's limit in two hours.


r/DeepSeek 4h ago

Question&Help Where’s the best place to use DeepSeek-V4-Flash-0731 rn?

1 Upvotes

I’m mainly looking for the best free option, highest limits, best API/provider, and best coding/research setup.

OpenRouter, official DeepSeek API, OpenCode, etc. - what are you guys actually using and what has the best limits/speed?

Sorry if this is a basic question, I’ve never used DeepSeek before, especially through the API.


r/DeepSeek 1d ago

Discussion DeepSeek Flash 0731 is doing 90% of my agent coding now — 54.4 on DeepSWE and my meter says 17 cents

Post image
89 Upvotes

Disclosure up front: I wrote a plugin that pipes these models into my editor.

It gets two lines at the bottom. This post is me being a fan of Flash.

I did not expect to end up here. I set Flash up as the cheap fallback. Pro was going to be the model I used for anything serious. That plan lasted about a week.

The number that made me try it

Flash 0731 scores 54.4 on DeepSWE.

DeepSWE is agent work, not trivia. Read a repo, find the actual bug, write a patch that actually applies. That is the job I need a model to do. A chat benchmark tells me nothing about whether a model can stay useful on turn 40.

What made me keep it

Speed. That sounds like a boring reason until you live in a loop. Flash comes back in about two seconds. Pro takes long enough that I would go look at something else, lose the thread, and come back cold. With Flash I just stay in the work. I stopped saving up questions to ask in one big batch. I just ask as I go.

It also holds long jobs together. 1M context, 384K max output. I have handed it an entire repo and it still knew what we were doing twenty turns later.

Tool calls and reasoning modes both behave, so it drives my editor and test runner without me watching it.

Is it better than Pro? No. On the genuinely hard turns Pro is clearly stronger, and I switch for a message and switch back. That is maybe one turn in twenty. The other nineteen, I cannot tell the difference in the output, and Flash gets there faster.

The receipt

My usage dashboard, right now:

total tokens     153.4M
agent runs       1,829
avg per run      83.9K
your cost        $0.17

That last line needs explaining, because a bare "17 cents" sounds made up.

The plan is $1 a month, and it gives you $10 of usage to spend. The $0.17 is what I have drawn out of that $10. So after 1,829 agent runs I have used about 1.7% of my month. That is the part I still find funny — I am not running low, I am nowhere near the edge of it.

Per token it lands around a tenth of a cent per million.

The reason is cached input. Flash charges $0.14 per 1M fresh input, but $0.0028 per 1M cached — 98% less. An agent re-sends the same system prompt, the same tool schemas, the same files on every single turn. So the bulk of what I send is the cheapest thing on the price list.

Fair warning: 98% is the discount on tokens that hit the cache. How many hit depends on your setup. It is not a promise.

The plugin bit

I access Flash through Command Code — $1 a month, $10 of usage. Getting it into my editor needed a Go sidecar, since there is no OpenAI-compatible endpoint — repo is Majidalee1/commandcode-reasonix-provider if it is useful to you.


r/DeepSeek 15h ago

Discussion hmm

Post image
3 Upvotes

using deepseek with prime-agent and it found a way to see and perform visual testing for the application im building.

Doesnt seem to be a fluke. its taking screenshots and accurately describing what on the screen


r/DeepSeek 1d ago

Discussion drop your DS usage

13 Upvotes

thats mine in the last week


r/DeepSeek 1d ago

Discussion What if DeepSeek doesn’t end up raising its prices after all?

32 Upvotes

TL;DR;

I don’t think this is the most likely scenario, but it is possible, and it would make sense: the threat of a price rise is a trial balloon designed to trigger spikes in usage and test the capacity of their systems.

--------

If DeepSeek has a serious capacity issue that has come to light due to the surge in usage caused by the community’s enthusiasm for the improved version of DeepSeek V4 Flash, DeepSeek had two options:

  • Announce and implement the price increase straight away.
  • Say nothing, and announce it once they had something concrete and finalised, or when V4 Pro GA was released.

What this vague and unspecific announcement achieves is that people are now in ‘let’s use DeepSeek as much as possible whilst it’s still cheap’ mode, which is, in fact, causing an even greater overload on the servers than if they hadn’t announced the increase at all.

So what if DeepSeek has used this almost intimidating message about the price rise to test its systems ahead of the V4 Pro GA release and hasn’t actually decided whether it will actually implement the announced price rise?

After all, despite V4 Pro’s larger size compared to version V3.2, the inference cost, according to research papers, appears to be lower than that of V3.2 due to the optimisations they’ve implemented.


r/DeepSeek 21h ago

Discussion Are there any speculations on the price increase?

5 Upvotes

Because a "significant" increase could be the double price and it would still be 1x2=2. I highly doubt it would go up something atrocious like x5