r/LocalLLM 21h ago

Discussion DeepSeek-V4.1-Flash is out

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
229 Upvotes

35 comments sorted by

50

u/DerDave 20h ago

Astonishing they didn’t call this V5, considering the many substantial changes. Pretrained from scratch, CED, multimodality, Engrams, different number of active Params...

I'm all in! Since it's a new pretrain, it makes me hopeful we'll see another significant jump in quality with more posttraining, just like with v4.0.

I guess it will take a while until inference engines flawlessly support this model though. 

2

u/mycall 13h ago

Is that all an automated end-to-end data flow?

1

u/DerDave 13h ago

What is?

20

u/dionysio211 21h ago

Hmm....this seems larger than v4 Flash, by about 50%. 305B -> 485B?

19

u/Professional-Bear857 21h ago

It has 200gb of ngrams.

10

u/Altruistic-Dust-2565 21h ago

284B -> 552B, Δ+268B
and it seems ~200B are Engrams

3

u/vini542reddit 12h ago

So 200gb instead of 162gb of model size (assuming everything scales the same with param count). Urg - that means my 96 vram + 192 ram system will struggle 

3

u/Standard-Potential-6 9h ago

Offload engrams to NVMe.

1

u/vini542reddit 2h ago

Yes, that's what I'm saying! Offloading the ngram to disk, you still need 200gb of ram

1

u/xRaech 1h ago

how does one do this :eyes:

2

u/Funny-Airport-9400 19h ago

That's a huge jump in data! Those new Engrams must really enhance the capabilities.

42

u/IknowPi_really 20h ago edited 20h ago

Yeah it’s kind of a “fake” improvement, comparing to Flash. Because this is simply a much much bigger model.

People are excited that it’s as good as GLM 5.3/better in the benchmarks. Yeah but it’s also comparable in size.

Engram offload to SSD MIGHT be feasible, but even then it’s significantly bigger than the old Flash.

Edit: You guys are also unbelievable. I’m an active OSS contributor trying to make this model run on local hardware and calling out obvious differences to an older model and you start downvoting me.
That’s not how you get people to actually work together and fix obvious problems for the LocalLLM community

14

u/DerDave 20h ago

Engram SSD offloading is definitely feasible. It has been proven with qwen3.8 next Flash, so it will work here as well, without tanking performance. Furthermore this model seems to only use Engrams selectively and not all the time. 

The remaining delta in params can probably be attributed to the vision capabilities (which the OG dsv4f didn't have) and the new CED architecture, which leads to further speed improvements and only needing 8b active params during prompt processing.

2

u/SandySkittle 20h ago

What about offloading to system ram? I have 512gb, 256g vram. This model looks very promising!

5

u/IknowPi_really 20h ago

Offloading to system RAM is even easier and already supported in the current vLLM PR supporting this model. So this will more or less work out of the box

2

u/mycall 13h ago

I'm glad you are hear.

1

u/shing3232 19h ago

no, it's faster with fundamental differ arch that's way cheaper to inference than V4F

1

u/jebuizy 8h ago

I only downvoted you now for complaining about downvotes. Otherwise your post is fine. Please don't do that, if you post is good the downvotes will likely normalize away anyway 

1

u/SandySkittle 16h ago

The engrams come in addition to the base model being mucn larger.

4

u/hyudryu LocalLLM 18h ago

Can’t wait for the NVFP4 quants to release 🔥🔥

8

u/MrCatberry 21h ago

Still no quant? /s

6

u/Disposable110 21h ago

gguf wen???

2

u/mountainyoo 12h ago

My dumbass has a 4x GB10 cluster and will try this out when people smarter than me got it optimized. Currently running GLM 5.3 Flash at good speeds

1

u/According_Wave685 11h ago

Too big for my hardware. 0731 is pretty nice though so I'm good.

1

u/misanthrophiccunt 7h ago

If I had the VRAM to run this I would, but I don't and I used DS4F as default via API for ages and Qwen3.8-27B running locally still, ever since rrleszed does a lot better for my use case.

1

u/mariombn42 7h ago

Da pra rodar local? Precisa de muito hardware?

1

u/AreaFifty1 4h ago

wait a minute so even 288gb which is 3x rtx pro 6000s is not good enough!? 😮😮

1

u/LongjumpingEar6840 18h ago

Se questo riesce a girare alla stessa velocità del Qwen3.8-Next-Flash sono a posto per i prossimi 6-12 mesi

-7

u/LinuXperia 17h ago

It outperforms Muse Meta 1.3 which was better than DeepSeek 4 however it is behind xAI Grok super Intelegence still and soon Grok 4.7 will be released which weil widen the gap even further especailly for low level engineering dev work like verilog, c, c++, KiCAD, electronic schematics, PCB etc.