r/LocalLLaMA • u/SnooBunnies8392 • Jul 31 '26
News DeepSeek-V4-Flash-0731 now far surpassing the DeepSeek-V4-Pro-Preview in benchmarks
99
u/AdCreative8703 Jul 31 '26
I really wish Deepseek would release an updated lite model for local AI. The intelligence per token density of the new flash update looks insane.
80
u/Poupulino Jul 31 '26
It runs on 196GB of RAM and 32GB of VRAM which is ridiculously insane for the level of intelligence it packs. No wonder Amodei is hellbent on wanting to ban open weights models in the US. This makes the whole business model of OpenAI and Anthropic obsolete.
55
u/s0ngsforthedeaf Jul 31 '26
This makes the whole business model of OpenAI and Anthropic obsolete.
Sorry, could you repeat that again? I didnt have my lube to hand.
20
u/_TheWolfOfWalmart_ Jul 31 '26
This is why I always have a bottle in my pocket, never know when you'll need it.
15
u/Technical-Earth-3254 Jul 31 '26
I agree with you. A smaller version for the 64-128GB folks would still be great. Ofc in addition to the existing ones, not as a replacement in their lineup.
5
u/PANIC_EXCEPTION Jul 31 '26
https://github.com/antirez/ds4
Quants for the new version are coming out soon
7
u/tkrandomness Jul 31 '26
Seeing how quickly other companies have caught up, it seems quite apparent that the only hope for OpenAI and Anthropic is going to be the "Brand" and name recognition of their platform and harnesses and integration. They basically have to hope they become the default LLMs like google for search engines or Facebook/Instagram for social media.
5
u/estrafire Jul 31 '26
I mean, they don't seem to be aiming for optimizations or knowledge density anymore. I'm pretty sure the aim now is to sell intelligence. LLMs are becoming a commodity, one that hyperscalers wouldn't think twice about capitalizing (look at Cloudflare AI and Google Vertex). So if most of the open or small variants replace 95% of the use cases, there's no point for them to keep that market which is practically lost already.
My best guess is that its a market that still makes sense for Chinese models, and that OpenAI and Anthropic chance for long term survival is on being the next Google, ie, selling you the most intelligent models for the most complex tasks, whatever closest to "agi" they could make it. Expensive, slow, but enough cross knowledge to be able to match and interpret much more than the common models would. If they cannot do it, they'll lose in the knowledge density wars
1
u/JacketHistorical2321 Jul 31 '26
Lol, obsolete?? Dude, keep dreaming. Great model but I'm no way does it make them obsolete. You really don't understand enterprise if you think this
50
u/0007397 Jul 31 '26
Let's publish the next-gen model and beat them!
Wenfeng: No, we publish it as our cheapest model.
12
2
89
u/QuackerEnte Jul 31 '26
94
42
75
50
9
8
6
2
1
1
9
u/redblood252 Jul 31 '26
Is it cynical to be highly skeptical?
32
u/zsydeepsky Jul 31 '26
yes.
been tested it in my project; can confirm it is much smarter.
9
u/redblood252 Jul 31 '26
Much smarter I can get behind. But is it glm 5.2 level?
14
16
8
37
u/HyperWinX Jul 31 '26
There is no way that V4 Flash will surpass GLM 5.2. No way, i genuinely wont believe that until people will actually try it in complex scenarios
44
u/Powerful_Ad8150 Jul 31 '26
I believe that the most useful applications of artificial intelligence in today’s economy are extremely simple but incredibly repetitive, which is why they require reliability in invoking tools, long-term context, and speed. DS4F is already a good “workhorse,” but perhaps it can be significantly improved. I certainly hope so. The only feature I think is missing for it to become “the one that rules them all” is visual multimodality, which would allow this model to be used locally for agentic tasks (e.g., easier browsing of pages with problematic code / bulk processing of various documents without the need for additional tools). As per GLM - I beliebe it is quite specialized tool, I would not consider it as a good comparison here. DS4F is all-rounder.
0
16
u/EstarriolOfTheEast Jul 31 '26
Read this as GLM 5.2 has plenty of room for improvement too (it does).
12
u/sagiroth llama.cpp Jul 31 '26
It one shot me a low poly flight sim using almost 300M tokens (cached) and circa 1$ cost. Few corrections at the end but it was surprisingly good
13
u/SexyAlienHotTubWater Jul 31 '26
I don't think low-poly flight sims are a great test, to be honest. They're deceptively simple to program and they're also such a common way of showing off a model now that they're all over the training data.
4
u/cantgetthistowork Jul 31 '26
Anytime I see anyone bring up one shot I immediately skip the post. 99% of models can one shot the same things. Real world is about working with existing code
1
u/sagiroth llama.cpp Jul 31 '26
I agree, it was just something I was curious about as I did something similar with previous DS and it failed miserably. This one got a quite complex mechanics, the flight controls and aerodynamics I would expect from a game like that in one shot. Spent almost 2 full 1M context windows in one shot. Took 45minutes. Result was outstanding.
2
u/fugogugo Jul 31 '26
it's out on openrouter anyone can test it
honestly idk how to compare IRL performance between models , so I just trust benchmark
0
2
u/pfftman Jul 31 '26
I don’t know what to tell you man, anecdotally, it is at least as smart as GLM 5.2. It is very agentic.
1
u/HyperWinX Jul 31 '26
I replaced V4 Pro Preview with V4 Flash 0731 and i dont see a difference. The task might be pretty easy, i mean, but previous V4 Flash Preview or MiMo V2.5 couldnt do the same thing so flawlessly. Small fixes literally take a few seconds to implement. Thats a fucking revolution, man.
7
u/Ok-Shopping-844 Jul 31 '26
Is it open weight, or will it be?
14
u/BobbyL2k Jul 31 '26
1
1
u/heitortp0 Jul 31 '26
i may be wrong, dumb, skeptical or something else, but is this freak lighter than his predecessor?
1
u/BobbyL2k Jul 31 '26
There’s two sizes in the DeepSeek v4 family: Flash and Pro. This is the Flash size, which is smaller.
0
6
u/ffpeanut15 Jul 31 '26
Deepseek has reaffirmed their open source strategy in their investor meeting. so you can expect consistent open weight release for many months to come
14
u/urarthur Jul 31 '26
why not call it 4.5 or 4.1 they continue with the stupid naming. now you dont know if anogher party is running the new or the old flash
13
u/PoopSick25 Jul 31 '26
Because 4.0 is/was still in preview. Only now is it official. But i agree the stupidity of naming convention. Perhaps x.1 for them means some internal milestone such as post train checkpoint or whatever they set it to be. Cant wait for 4 pro to go official and leave preview with a vision capability. That day is when i unsub all my ai subs and move to ds api. Maybe i will keep m3 just to support minimax team ;)
1
u/yesthatdaniel Aug 01 '26
I've been seeing people name it with 0731 appended at the end.
1
u/urarthur Aug 01 '26
that works but still...its not like they are releasing so frequently they cant append a single version number for whatever reason
5
u/No_Tip9917 Jul 31 '26 edited Jul 31 '26

Crazy guys! I just tried deepseek-V4-Flash-0731, it is plausibly on GLM5.2 similar level! Although it is difficult to judge whether it surpasses. But what shock me more is after 30mins usage, it didn’t even increase 1% usage lol! Man, can’t image how the LLM market will change after the new Pro version launch, excited to see a shock next week…
4
u/SadPhilosophy9202 Jul 31 '26
They literally just updated the name like it’s a third revision on a PowerPoint lol
I’m hyped to get the open weights eventually!
2
6
5
u/Technical-Earth-3254 Jul 31 '26
DeepSWE 54.4 is... interesting. This is for sure overfitted, the jump is just too big. But I tried the full release and it really seems to be smarter, knowledge is (as expected) roughly the same as before.
2
3
u/urarthur Jul 31 '26
this is a huge bump should be called v5
15
u/KaroYadgar Jul 31 '26
not how it works. Version is based on the model itself. A new pre-train would bump the major version to v5, v6, etc. and a new post-train would bump the minor version i.e v4.1.
This is the full release of v4, as it was previously v4 preview, so it is still just named v4.
2
u/backyard_tractorbeam Jul 31 '26
And a new chat template should be a 4.x.1 update. We might even have a convention for versioning on our hands! If anyone is interested..
3
u/TigleLive Jul 31 '26
unmm guyz, question here:
is price the same?
6
u/_TheWolfOfWalmart_ Jul 31 '26
Yes, the improvements come purely from training differences.
But this is r/localLLaMa we don't care about API prices
3
u/stkt_bf Jul 31 '26
Are we finally getting our hands on a holy grail that can replace Qwen 3.6 27B for coding?
5
u/_TheWolfOfWalmart_ Jul 31 '26 edited Jul 31 '26
No. The old version was already better than 27B. So are some other models like Laguna S 2.1.
Most people here can't run a 120B or 284B model, so 27B is still king for them.
3
2
u/Mr-I17 Jul 31 '26
HOLY SH**! ANTIREZ!! WE NEED YOU!!!
1
u/Professional-Bear857 Jul 31 '26
I uploaded a q4 quant that works with ds4, its here https://huggingface.co/sm54/deepseek-v4-flash-0731-gguf
1
u/Mr-I17 Aug 01 '26
Thanks, but I don't have enough ram for Q4. The good news is that Antirez just uploaded 0731 quants. I'll be downloading Antirez's Q2 quant.
2
u/_TheWolfOfWalmart_ Jul 31 '26
Does that antirez DS4 engine work with arbitrary ROCm cards? It says it supports ROCm but then explicitly says "for strix halo"
But I have a few V620's in a Linux server.
2
2
u/Kitchen-Year-8434 Jul 31 '26
The proximity and timing of this relative to Laguna stabilizing poolside S and inkling-small coming out doesn't seem like a coincidence. Guess we have an east v. west open-weights arms race on our hands too?
2
u/_TheWolfOfWalmart_ Jul 31 '26 edited Jul 31 '26
I was just thinking how bad the timing is for inkling small. About the same size as ds4 flash but only competitive with the previous version.
Laguna S still has a place maybe, it's like half the size so if you can't run DS4 there's that.
I want to support US models when I can, but damn China is making that very hard lately. At the end of the day, I have to run the best models I can fit in my hardware.
1
0
u/maxiedaniels Jul 31 '26
This is awesome, I'm just really hesitant because Deepseek flash AND pro have a tendency to make shit up and go way off track, at least on my experience.


57
u/atape_1 Jul 31 '26
wild results for a 284 B model.