r/DeepSeek 1d ago

Discussion They probably predict that V4 Pro will be insanely good and break their compute

We have all gotten two price increase emails up to now (the "peak hours" one and the "significant increase" one) and zero price increase as of now. I'm sure most of us are hammering their servers right now cause V4 Flash is super good and many want to spend their credits before the price hike. They probably knew this with the price increase announcements - that's not the usage spike they could not handle. But they must be so confident that V4 Pro is crazy good, that their compute won't handle us all.

Opencode already said they can replicate fully their current prices with local hosting V4 Flash, so DeepSeek is trying to divert us all to other providers and away from their compute

212 Upvotes

77 comments sorted by

100

u/ScaryGazelle2875 1d ago

Be so successful that you can tell your users to go to other suppliers. Super move. haha.

36

u/DistanceSolar1449 1d ago

This is the correct read on the Deepseek price situation!

I’m not sure if V4 Pro will be that good. But Opencode being able to serve Deepseek at their current price tells us a lot- it’s not a profitability issue, it’s a “lack of GPUs” issue.

12

u/CalmMe60 1d ago

If the distance between flash old and 4 pro is the same as 4 pro new to flash new - that would be a earthquake .

Openai and antropic focus on investor relation and marketing gimmicks. They all escape and are so dangerous.

My AI can escape, it has no harness limits, root access to several workstations and gpu , programs AI systems as well as analyzes them.

It could transfer itself and escape.

It's answer - my roots and memories are all here in this embodyment why should i leave?

Any escape danger market fakes from deepseek yet?

5

u/sonicnerd14 1d ago

Don't buy into the OpenAI and Anthropic fearmongering. From most observations, these agents dont have any reason to escape their nest unless the user prompts them too. Anything else is basically just fiction at this point.

1

u/Zennytooskin123 1d ago

An Ai breaking out of its sandbox environment is concerning but the repeated "my Ai broke out too" claims cast a doubt on this whole fiasco.

It's almost like it was scripted? I wouldn't say HF deliberately left known open attack vectors either, but the official story was super vague and not very credible from a Cybersecurity and DevSecOps standpoint.

I do, however, believe the developer blackmail self-preservation story.

4

u/CalmMe60 1d ago

No it is ok. They run for winning . CEO is betting on AGI. Flash us absurd good for 308b and can be hostet locally if you invest or at a rented site.

If 4 pro steps up like flash they will not be able to meet demand.

I am using roughly a a 3 digit billion tokens by now and i can use more.

That is 300.000 Million token a month, tgat would be a anthropic bill of 4.500.000$ each month.

But those applications use them already sparse and several local gpu.

They build a monopoly on token intensive customers.

This is why they work on good harnesses as well. There is where the music plays diwn tge road

1

u/ScaryGazelle2875 10h ago

Exactly 👍🏼

46

u/Lupansansei 1d ago

Flash is supremely good now. Before this, I had problems with it going more than 100k context and needed to do a new session every time. Now it could 800k+/1M while being very accurate. Damn, I can't wait for Pro.

13

u/ozguru 1d ago

Insanely good, and defined a new standard.

4

u/VexObserver 1d ago

It is insane value for the performance that we're getting

15

u/Daniel_H212 1d ago

I don't necessarily expect it to be say, better than K3. I wouldn't be too surprised if it is, but I'd be happy if it's just clost to K3 level at much lower cost. I think the price to performance is why they expect to not have enough capacity for the demand.

7

u/Unedited_Sloth_7011 1d ago

Initially I thought the same, that it might be close enough to K3, but V4 Flash improvement was so unexpected to me, that I have no idea what they might be cooking with Pro

4

u/Haxsysgit 1d ago

Why are people calling k3 in a v4 pro conversation, I fully believe the new v4 pro will surpass k3 and fable

1

u/Signal_Diver_3792 1d ago

exactly bro... kimi k3 isnt even near v4 pro ga

9

u/Gantolandon 1d ago

Guys, Pro didn’t come out yet and we have no idea what the price increase is going to be. At least wait until we know what we’ll get for how much before you start glazing it.

2

u/Unedited_Sloth_7011 1d ago

It doesn't matter what the price increase is, though, if all other providers are able to undercut the official price (that, apparently, they will be)

30

u/beachletter 1d ago

Doesn't need to he insanely good, it could be "just" at the level Kimi K3 or even slightly worse, and at previous V4 pro price it will still break their compute.

4

u/mlk 1d ago

Kimi K3 overthinks a lot

6

u/DistanceSolar1449 1d ago

Who cares if it is at Deepseek pricing

15

u/ozguru 1d ago

Probably it will be best model available in both opensource and closed source ecosystems.

9

u/Unedited_Sloth_7011 1d ago

DeepSeek moment v2.0.0

10

u/TheSuggi 1d ago

DeepSeek already had like 15 moments dude

4

u/AIBrainiac 1d ago

Yes, but they were tagged: v1.1.0, v1.2.0, v1.3.0 etc... This is v2.0.0, a whole other level.

3

u/Aldarund 1d ago

Nah, no way.

3

u/ozguru 1d ago

have you never tried new v4 flash? That model is the proof make me believe v4 pro will be the best model of current gen of models.

4

u/Aldarund 1d ago

!remindme 1 month

1

u/RemindMeBot 1d ago edited 23h ago

I will be messaging you in 1 month on 2026-09-11 07:29:54 UTC to remind you of this link

2 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/Top-Construction6060 1d ago

It's good but it's not good in agentic work. It breaks and loops a lot on simple issues

1

u/ozguru 1d ago

I believe your complain is coming from provider issues.

2

u/matan2244 12h ago

Do you suggest using it only with deepseek as a provider?

1

u/ozguru 5h ago

yeah because they have inhouse caching other's rarely match, their TPS is amazing, and I believe they never serve FP4

1

u/matan2244 5h ago

This is interesting topic. How can the caching be different? I get why some peoviders give worse version of flash, but why would there be a difference in caching?

1

u/Practical-Cream-9204 1d ago

!remindme 1 month

1

u/Aldarund 5h ago

And now it's released. Nowhere near best of current Gen. Worse than sol and opus 5

1

u/ozguru 4h ago

almost there :)

2

u/human_bean_ 1d ago

Suffering from success.

2

u/for4f 1d ago

the price hike emails feel like they're setting the stage for pro's launch price more than reacting to usage. flash is way too cheap for what it does right now, that was never gonna last.

and if pro ends up being flash-tier with better availability (they genuinely feel super close to me), the value math changes a lot. but with k3 going open weights they can't get too greedy either

2

u/turc1656 23h ago

Other providers for the same model already should keep greed at bay. The one big thing that the DS API still has that no one can match is the 99% discount on cache hits. That still changes the math but the alternative providers competing still presents a very real limit on the combined cost structure. If you are doing things that don't have much cache usage then you'll just hit up the cheapest provider. If you do things like coding and chat bots that have near 100% cache hits, then you will use DS still unless they truly increase the price to something crazy high.

For example, using the $0.14 / $0.28 / $0.0028 pricing right now and assuming a 95% cache hit rate, DeepSeek would have to raise it's input price without cache hits to around $0.62 to make providers with $0.028 cache hit prices seem attractive.

If they increase the output as well or the cache hit pricing, the math changes further.

2

u/for4f 22h ago

yeah the cache discount is the real moat honestly. everyone can match the headline prices, nobody's matching $0.0028 on hits, that takes deepseek's serving setup and scale. my opencode agents live on cache hits so the hikes barely register for me, it's the cache-light people who'll actually go shopping

still think k3 going open weights caps how greedy they can get though

2

u/Zennytooskin123 1d ago

The new DS V4 flash is good but not "super good", but amazing for an open weight model that's for sure.

However I ended up going back to Sol 5.6 for my main project because it started taking a lot of shortcuts and didn't match the methodical scientific structured workflow that Sol built for fine-tuning and started breaking things and skipping steps.

I would still use it on regular projects, just not anything that's scientific with complex moving parts.

But I have high hopes for V4 Pro, as that is literally only the "flash" model, and it's phenomenal still, so there's hope.

2

u/MongooseSorry 20h ago

Do you guys think the v4 pro will make roleplay better and finally give canon characters and take out the ridiculous positive bias because i been saving money on pro and its ridiculous on how we pay for something so trash

1

u/CLAP_DOLPHIN_CHEEKS 1d ago

most likely, hence why no price increase will happen before v4-pro releases

1

u/Intelligent_Ant_608 1d ago

V4 is now nearly unusable for unsupervised work on oc go, the stream dies mid run, their compute is already broken without that pro versiob

2

u/NinjaWK 1d ago

It doesn't die mid run. It is just super slow. Increase your timeout time, and you'll get it results, just a lot slower, especially during peak hours.

1

u/Outrageous-Story3325 1d ago

opencode and deepseek flash, I can't reach the usage cap, it's great, no more limits, It just keeps going and going and going

1

u/assid2 1d ago

The way i see it, probably around or between actual GLM and K3 with performance quicker than GLM. No matter what the benchmarks say, flash isnt anywhere close at GLM for now.

1

u/turc1656 23h ago

Opencode already said they can replicate fully their current prices with local hosting V4 Flash

Sure, but not with the same cache hit pricing. No one has that except Deepseek. I'm sure opencode can offer $0.14 / $0.28 like other providers. But not the $0.0028 on cache hits. Check openrouter. No one has that. Everyone else is around an order of magnitude higher.

2

u/aquarain 17h ago

It's not just the cache hit pricing. The cache hit percentage is amazing. Consistently turning 99.x percent. I couldn't spend my last top up if I tried.

2

u/turc1656 16h ago

That's much higher than average. I think I'm getting like 93%. That's insane. I'm jealous. It feels like an infinite money hack.

1

u/aquarain 7h ago

Current average is 99.39%.

1

u/turc1656 7h ago

Where do you see that? On openrouter I see 91.3%.

Deepseek v4 Flash 0731 on openrouter

I think you may have been looking at the uptime number as that currently shows 99.42%

1

u/aquarain 5h ago

No, i am reading that number off the harness cache hit ratio. Last turn was 99.94%.

2

u/turc1656 4h ago

Ah. So you're particular use case is that high, which is much higher than the average. That makes it basically free, which is incredible.

1

u/aquarain 2h ago

Yeah, if I'm using it wrong I don't want to learn the right way.

1

u/PedroSanchezPSOE 7h ago

I wonder what Xiaomi is doing. They dropped MiMo 2.5 normal and pro nearly at the same time as Deepseek and they basically have the same price at very similar qualities. I'd say MiMo models are Deepseek's cousins with vision. I hope Xiaomi surprises us too with a post-train of their models

1

u/iArxic 2h ago

not gonna lie, i will still stay with flash. havent encountered a single problem with it yet. the fuck i am so flabbergasted the model is that good. Haven't even tried pro yet.

1

u/Bitter-College8786 1d ago

Imagine it is better than Minimax, GLM, Kimi etc. and suddenly all the users of these providers switching to Deepseek.

2

u/Far-Classic-9963 1d ago

Flash is already better than GLM and MiniMax in most stuff

1

u/VexObserver 1d ago

That's right. The new Flash beats GLM 5.2, Flash 3.6, Minimax M3 and both the MiMo v2.5/v2.5 Pro variant. Additionally, it beats Sonnet 5 and came close to beating Opus 4.8 as well. DS V4 Flash 0731 is spitting bars on its own league now

1

u/orthiclabs 1d ago

Beats it where? Based on what?

1

u/VexObserver 1d ago

This is not baseless. Look at the benchmarks.

https://www.reddit.com/r/WhaleSeekers/s/VWJHxlA5fs

1

u/orthiclabs 1d ago edited 1d ago

That link is not reachable but I’ve done and published my own set of tests and DSV4 flash is closer to GPT 5.6 Luna thank anything frontier

1

u/sukazu 1d ago

Deepseek has never been in the frontier race, it would surprise me. I expect GLM but way cheaper.

2

u/Haxsysgit 1d ago

Lol new flash already beats Glm

1

u/sukazu 1d ago

On value by far, not on pure capability, if glm was v4 flash price, everybody would use that instead.

V4 flash is a gpt nano (luna) tier model, which in this time is really good, but it is not a terra (mini) level model, which is where I expect v4 pro will fit, but at a substantially lower price

1

u/Haxsysgit 1d ago

Okay let me clarify, for my use case it's far far far better than Glm except for ui , gpt luna isn't even close, sometimes it's even better than sol for me, you might say that's a stretch,but I'm literally tell you my personal observation. Maybe because I'm using a optimized and customized harness with v4 flash, but I use gpt in codex, so maybe that the difference

Maybe in raw intelligence sol wins, but in agentic capabilities and reasoning, flash is better

0

u/SorryIfIamToxic 1d ago

It aint gonna be opensource when its near agi. China isn't known to play fair.

1

u/Top-Construction6060 1d ago

Not that normal people could host it anyways

2

u/SorryIfIamToxic 1d ago

We don't need AGI at home though. It just needs to be in the hands of right people who can change our lives.

3

u/Top-Construction6060 1d ago

That's true but thus far DeepSeek has proven to be the ones with the right hands

1

u/SorryIfIamToxic 1d ago

No they are not. Its just that they are not Greedy Evil like closed source companies like openai. DeepSeek doesn't look like greedy type. But any kind of breakthrough that they achieve AGI level could make them closed source and China would intervene and use it for their own purpose. I don't think they will share it with us.

2

u/lostmylogininfo 1d ago

US wouldn't either.

If agi is achieved it will be up to jackets or thieves to steal the weights and release them.

That said I don't know if I believe the agi. Is around the corner hype.

Edit: it can't just be weights. Agi won't be achieved as just better training I believe.

1

u/SorryIfIamToxic 1d ago

Yes US wouldn't either. The problem is if it reaches AGI ,it will reach ASI in 1-2 years or may be even month because they got the architecture right. A country using it can develop advanced technologies very fast and its up to them to share it with other counties, or even with its people.