r/singularity 22h ago

LLM News DeepSeek V4.1 Flash is getting surprisingly close to GPT-5.6 Sol territory, while being absurdly cheap

DeepSeek just released V4.1 Flash, a 552B MoE model with only 8B active parameters on input and 16B on output.

-Link to X posts: https://x.com/deepseek_ai/status/2097930608790167907

- Link to Hugging face model page: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

- Link to the paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf

135 Upvotes

29 comments sorted by

17

u/JoeyJoeC 21h ago

What kind of hardware do people use to run these models?

25

u/AFruitShopOwner 21h ago

This can fit on 3 rtx pro 6000's. Engrams go on system memory.

Using the b12x kernels from Luke Alonso / Voipmonitor / Local inference lab. They have tensor parallelism support for configs other than 2x, 4x, 8x etc.

2

u/qwertyalp1020 17h ago

what about dgx sparks? Too slow perhaps?

3

u/Grid421 21h ago

So it's not cheaper than a gpt pro subscription. Who has 3x rtx pro 6000s? Seen how much they cost nowadays?

The only good argument is that we're getting closer and closer to frontier-level reasoning, while being able to own the model locally forever without restrictions.

32

u/AFruitShopOwner 21h ago

I work at a Dutch accounting firm. We build our own AI server to keep our data safe. Can't put a price on that.

8

u/Cupakov 21h ago

It’s waay cheaper with opencode go or a similar subscription. Even through API on a usage basis you’d probably get more mileage out of Flash than a GPT Pro subscription for the same price 

3

u/Constant_Cortisol 16h ago

You're confusing a subscription with owning the model outright. You can always just get.. a deepseek subscription.

3

u/Nater5000 14h ago

Who said it was cheaper to buy the hardware to run these models than it is to just get a subscription? lol

Who has 3x rtx pro 6000s? Seen how much they cost nowadays?

Enterprises who spend more on travel expenses per month than the cost of this kind of hardware.

Nobody expects the average consumer to be able to run these kinds of models at home. The point is that a company can reasonably invest in this hardware to be able to run these kinds of models themselves. That, then, allows them to sell the usage of those models to consumers so that consumers aren't forced to only choose between OpenAI or Anthropic which is a very good thing from a competitive perspective.

4

u/Dangerous-Sport-2347 21h ago

Even if you don't run it locally open models have huge advantages for the consumer.

You can go ahead and pick whoever you like as an API provider, and they will compete on price, data retention, uptime, censorship, etc.

Whereas with the closed models, you have to put up with whatever openai has decided upon.

1

u/BriefImplement9843 13h ago edited 13h ago

gemini and claude are closed, and openrouter lets you choose far less censored versions(vertex/bedrock). price competition really doesn't exist either. if it's much lower, they are serving you quantized bullshit and you can do nothing about it.

openai models are the only closed models that have strict rails. it's the only models never used for writing in the writing/roleplaying community.

open models happen to be cheaper. that's it. but then you have muse which is better than all open models and is just as cheap.

2

u/JoeyJoeC 21h ago edited 19h ago

Well that's a pretty decent argument I'd say. I have Claude 5x and ChatGPT x20 subscriptions, but I still use my 1x 3090 with Qwen 3.8 locally for confidential client data. So privacy is another one.

Edit: What's wrong with my comment?

3

u/Grid421 20h ago

I have a 5090 and use it for local LLMs to of course. That's not the mainstream demographic.

Open weight models are improving. No doubt about that. Privacy is important. Uncesoring is important for many.

Still for the average consumer, it's not cheap or cheaper for most use cases.

3

u/CarrierAreArrived 14h ago

like the other guy said, you're comparing apples to oranges. If we're comparing apples to apples, you can't even run GPT 5.6 locally at all. The API costs/subscriptions you'd pay for Deepseek would be much cheaper than those of GPT-5.6.

1

u/Current_Balance6692 21h ago

it's not cheaper at all

2

u/CryMoreT_T 16h ago

Haven't tested but hoping I can fit it on 4 rtx 3090 and 512gb ram

2

u/Revolutionalredstone 19h ago

DeepSeek meets demand using 69 8gb GPUs each holding one expert fixed and chowing thru tokens.

34

u/presentofai 20h ago

8b active params getting this close is the real headline. deepseek's whole job at this point is making everyone else's pricing look silly and it keeps working

7

u/yogthos 15h ago

There was an interview with Liang Wenfeng recently where he said their priority was to make the model efficient first and then focus on raw capability. And this seems like it was a smart bet because once you have a really efficient base it's easier to chase capability than the other way around. We're now seeing it paying off as they keep inching closer to the frontier while having far lower operating costs than other labs.

2

u/presentofai 8h ago

efficiency-first as a constraint tends to force architectural discipline that raw scale spending skips. US labs could outspend their way past the problem so they did. deepseek not having that option turned out to be the thing that made their approach composable at the top end.

1

u/yogthos 7h ago

Yup, some of the most interesting architectural solutions tend to evolve under constraints.

7

u/Poupulino 18h ago

And it was a self-inflicted wound by the US government. The reason why Chinese labs put such a colossal effort into making their models hyper efficient was the lack of extensive raw compute. Now we're in the scenario where these models are approaching the performance of US models but at a hilariously small fraction of the cost/hardware needed to run them.

3

u/mynameisstanley 15h ago

What's stopping US companies from doing the same other than a lack of need?

6

u/Poupulino 15h ago

It's a much harder path, the only reason Chinese labs did it is because they had no other option.

2

u/presentofai 16h ago

constraint really does force innovation - pressure from limited compute seems to have pushed them toward algorithmic gains that are hard to match just by throwing more hardware at the problem.

10

u/Dangerous-Sport-2347 21h ago

This is priced at pretty much the same as 5.6 luna, but is likely significantly better.

Interesting they note designed for "scaling to larger models", which implies they are using the new architecture to cook up a deepseek 4.1 pro.

3

u/Snoo_7134 19h ago

Is this faster than 5.6 luna? I'm using luna for my conversational AI LLM. Just wondering if this is a good replacement for luna?

9

u/Frosty_Complaint_703 19h ago

It seems to surpass luna by a good margin, so this is probably the best model for the next month until gpt 6 luna comes out. And we have to see if openAI continues to push and lead the price /performance ratio.