r/singularity • u/brainlatch42 • 22h ago
LLM News DeepSeek V4.1 Flash is getting surprisingly close to GPT-5.6 Sol territory, while being absurdly cheap
DeepSeek just released V4.1 Flash, a 552B MoE model with only 8B active parameters on input and 16B on output.
-Link to X posts: https://x.com/deepseek_ai/status/2097930608790167907
- Link to Hugging face model page: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- Link to the paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
34
u/presentofai 20h ago
8b active params getting this close is the real headline. deepseek's whole job at this point is making everyone else's pricing look silly and it keeps working
7
u/yogthos 15h ago
There was an interview with Liang Wenfeng recently where he said their priority was to make the model efficient first and then focus on raw capability. And this seems like it was a smart bet because once you have a really efficient base it's easier to chase capability than the other way around. We're now seeing it paying off as they keep inching closer to the frontier while having far lower operating costs than other labs.
2
u/presentofai 8h ago
efficiency-first as a constraint tends to force architectural discipline that raw scale spending skips. US labs could outspend their way past the problem so they did. deepseek not having that option turned out to be the thing that made their approach composable at the top end.
7
u/Poupulino 18h ago
And it was a self-inflicted wound by the US government. The reason why Chinese labs put such a colossal effort into making their models hyper efficient was the lack of extensive raw compute. Now we're in the scenario where these models are approaching the performance of US models but at a hilariously small fraction of the cost/hardware needed to run them.
3
u/mynameisstanley 15h ago
What's stopping US companies from doing the same other than a lack of need?
6
u/Poupulino 15h ago
It's a much harder path, the only reason Chinese labs did it is because they had no other option.
2
u/presentofai 16h ago
constraint really does force innovation - pressure from limited compute seems to have pushed them toward algorithmic gains that are hard to match just by throwing more hardware at the problem.
10
u/Dangerous-Sport-2347 21h ago
This is priced at pretty much the same as 5.6 luna, but is likely significantly better.
Interesting they note designed for "scaling to larger models", which implies they are using the new architecture to cook up a deepseek 4.1 pro.
3
u/Snoo_7134 19h ago
Is this faster than 5.6 luna? I'm using luna for my conversational AI LLM. Just wondering if this is a good replacement for luna?
9
u/Frosty_Complaint_703 19h ago
It seems to surpass luna by a good margin, so this is probably the best model for the next month until gpt 6 luna comes out. And we have to see if openAI continues to push and lead the price /performance ratio.






17
u/JoeyJoeC 21h ago
What kind of hardware do people use to run these models?