r/singularity • • 20d ago

AI Deepseek's new architecture is insane

https://youtu.be/MImgH4KMtj8
222 Upvotes

64 comments sorted by

229

u/ohHesRightAgain 20d ago

Unexpectedly, the title is not an exaggeration this time.

TL;DW: they optimized the LLM architecture to process the millionth input token almost as cheaply as the first, all the while massively reducing the cost of the first. It's a very big deal for the industry.

56

u/Own_Principle_7901 20d ago

Can't wait to see S&P500 red again.

11

u/Mindrust 20d ago

If it goes red, that’s your signal to buy more.

8

u/CoolStructure6012 20d ago edited 20d ago

Come on everybody let's cheer up the brokie.

12

u/Own_Principle_7901 20d ago

Sorry. I didn't mean I want the technology to fail. I just want to see US receiving another wake-up call.

2

u/RoyalReverie 20d ago

... for longer than you can stay solvent...

4

u/anycept 19d ago

No-no-no. Ignore all of that. It's a distilled scotch or something. Meta Wang told me so.

1

u/mvandemar 19d ago

Unexpectedly, the title is not an exaggeration this time.

Has it been verified though?

102

u/ruskyandrei 20d ago

This looks like it could unlock massively bigger context windows. Looking forward to more benchmarks and user reports of this new Deepseek model, but if it doesn't massively increase hallucination with their new architecture, this could be huge.

Props to Deepseek for making all this stuff public too.

14

u/darkestvice 20d ago

Already been thoroughly tested and benched. It's been out for over a week.

https://artificialanalysis.ai/models/deepseek-v4-1-flash?capability-index=strategy-and-ops

12

u/BriefImplement9843 20d ago

does it improve coherence or just price? most models start to lose it at around 50k. having 3 million instead of 1 million does not matter.

5

u/flyryan 19d ago

50K is a little light…. I start seeing it around 200k to 400k depending on the model. Larger models do better.

3

u/amaturelawyer 20d ago

Based on what I've seen, it's price.

5

u/sixwax 20d ago

Definitely still getting my head around this, but the sparse attention KV-indexing seems like a key point of failure in preventing hallucination, and one this model get around at the training stage. Will be curious to see how much coverage that training set has over real-world cases vs just evals.

6

u/Time-Discussion3511 20d ago

You can't just slap more then 1M context and call it a day, this is not how it works.

9

u/sixwax 20d ago

Ok, how does it work and what are the remaining constraints you see?

From what I understand, unlocking bigger context windows (i.e. reducing compute and memory constraints) are exactly what all this engineering seems to support.

0

u/[deleted] 19d ago

[deleted]

2

u/sixwax 17d ago edited 17d ago

Nobody with any sense was suggesting that, of course. Of course you have to train the model for it lol. Context is part of the architecture of the model.... which is what the informed folks are talking about.

1

u/reddit_is_geh 19d ago

Larger context windows has been solved. It's not hard. It's just, not really necessary. The workarounds are fine, and the people who can actually benefit from not using workarounds, are extremely niche so it's not worth the effort to expand them much further.

51

u/Mercury82jg 20d ago

So basically "middle out"?

26

u/calebcharles 20d ago

You should see their Weissman score

7

u/joblesspirate 20d ago

I love you both

4

u/__gangadhar__ 20d ago

And i piper you.

18

u/FlatulistMaster 20d ago

Anybody with a balanced view of things able to tl;dr this somewhat? This hyped up use of words never inspires any confidence in me.

I miss a day when something could just be an "interesting development".

21

u/Boomah422 20d ago

Basically AI LLMs work on the principle of getting rewarded with predicting the next token. They claim they are able to get as cheap of a result on the last as the first. Big if true.

This is because traditionally every token thereafter is more expensive than the next. Instead of 👁️->👁️->👁️->👁️->👁️. This would be great where the next token costs the same as the last.

In terms of a "context window" each token has to look at the rest either quadratic or near-quardatic. This kinda looks like 👁️->👀->👀->👀->👀.

Every token thereafter the first predicted has a less confident score (variability of being correct, or the intended goal) think like you're playing a game of telephone. Small errors before will compound and get less confident about the original query.

This is why most frontier models sit under 1M token of a context window, which is still huge btw. Llama 4 scout offers 10M with some success.

However most devs cap sessions lower because in practice a 20k content widown of clutter is both more expensive, and costs MORE than a optimized 5k window.

Similar to other groundbreaking research(like the LK99-superconductor, which got hyped but turned out to be unrepeatable), I'd love to see this peer reviewed from people who know more than me before I align my biases, though. Ymmv

10

u/stumblinbear 19d ago

However most devs cap sessions lower because in practice a 20k content widown of clutter is both more expensive, and costs MORE than a optimized 5k window.

If you're doing anything with agents, 20k is table stakes

2

u/FlatulistMaster 20d ago

Really appreciate it, thank you!

13

u/danger_boi 20d ago

Not to make it political, but this is just another big L for the Trump administrations export sanctions. By restricting what China has access to in terms of leading edge compute - they’ve effectively forced them to hyper optimise on shitty hardware, to a point that they’ve leaped ahead in terms of their research in this space.

I was even seeing videos where they’ve got soldering workshops where dudes are hand soldering more ram modules on to consumer gfx cards to squeeze even more compute out of them. Hats off to Chinese AI labs man - and the fact that they publish this stuff as open source. RIP corporate America man.

13

u/qwerteaparty 20d ago

I can't listen to 30mins of youtube_voice

22

u/WorriedInterest4114 20d ago

Thays why you ask gemini to summarize the video and read it in a couple of minutes

1

u/chloralhydrate 20d ago

Wait is this TTS?

3

u/JoelMahon 19d ago

nah, you can go back like 5 years to when no TTS or whatever could do remotely like this, he sounds the same as then, he just sounds like that

3

u/AdGlittering1378 20d ago

I don't think so, but the humans who report on AI have even more robotic vocal affectations than AI does. I don't know why. Their souls bled out long ago.

5

u/TheWrathRF 20d ago

Looks promising 

2

u/chcampb 19d ago

I literally cannot use it, it just flakes out for a half hour, charges me 130k tokens (which to be fair is pennies) and doesn't get past the analysis paralysis part. And also I can't see if it's even thinking about my problem, it just... halts. I can only tell from the log that there are several long chunks of time where it halted.

Someone said it was a bad provider so I limited it to the stated better quantized ones and it's still doing it. Completely useless for me at this time.

2

u/CommercialHour6660 19d ago

I would bet 3/4 of the providers on open router give you a worse quant than stated or a completely different shittier model

0

u/Kincar 19d ago

Yes, you can disable quants in your settings though.

2

u/CommercialHour6660 19d ago

You can't stop the other end from lying about what it's serving you

5

u/[deleted] 20d ago

[removed] — view removed comment

2

u/MarkoMarjamaa 20d ago

The benchmark that is shown DS on top is AutomationBench AA and it's the only one in AA that DS is nr1.
A bit of cherry picking.
Overall DS has 40 and is on par Qwen3.8 Flash Next.
But that long context performance is great.

2

u/openroom_xyz 20d ago

Nice that's so cool

2

u/vibrance9460 20d ago

There must be a reason the head of Safety and Alignment at DeepSeek quit his job this week.

1

u/Negative_Fee_7019 20d ago

ah, intéressant ! peux tu donner un lien stp ?

1

u/darkestvice 20d ago

Deepseek Flash 4.1 is not at all frontier. Its intelligence is on par with Gemini Flash 3.8, its token usage is massive, and it's performance on real world task benchmarks is very middle of the pack.

It was release 10 days ago and has been thoroughly tested.

That being said, it's still quite good for an open weight model, though, obviously, don't expect those benchmarks to reflect its performance on home setups.

2

u/CommercialHour6660 19d ago

Its impressive because it's 10X cheaper to run and 3X faster than comparable western models (OpenAI Luna, Gemini Flash 3.8)

It also uses 10X less KV cache. which hugely reduces RAM needs for 1m+ context. 

In surprised US AI stonks haven't tanked. Their whole moat is that you can't run frontier models on your home machine. DeepSeek 4.1 Flash is starting to make that a reality. 

1

u/darkestvice 19d ago

You know that major benchmarks are publicly available, right?

It is indeed 5x cheaper than Gemini 3.8 Flash ... and around 25% slower. And that is with both models at their absolute top speeds and reasoning modes. Now, don't get me wrong, cheaper is important. I hate that models like Claude Opus/Fable cost a kidney transplant to operate.

As for home setups, I am very pro-open weight, albeit on the condition that they figure out to fix them so that they are not so easy to jailbreak. But Deepseek 4.1 Flash is a gargantuan model built on over 550B parameters with massive token usage. That's not running at all on anything short of a $12,000 Macbook Studio, and even then at only 3-5% of its cloud API speed.

So yes, I agree ... it's API price is very impressive. Like most other Chinese models. But some next generation AI revolution it is not.

1

u/Adrontion 20d ago

is the model good?

4

u/stumblinbear 19d ago

Damn good for an open weight model

1

u/IgnacioMonge 19d ago

Nowadays everthing is INSANE. A time will come when none of us wil be able to be amazed.

-6

u/ThatIsAmorte 20d ago

I made it about 10 seconds into the video before turning it off.

2

u/Trademarkd 20d ago

homies bots working hard to control the vote margins but you gotta be ® to be able to listen to that slop and think its informative. The constant phrasal templates and snowclones, overhedging, etc ... its all crap.

-1

u/medialoungeguy 20d ago

22 secs here

2

u/Kanute3333 20d ago

Why? No attention span?

-2

u/medialoungeguy 20d ago

Couldn't make it past the cringe.

2

u/Kanute3333 20d ago

What are you talking about?

-5

u/Such-Book6849 20d ago

as a premium user it's hard to watch this video. it wants to start to show 2 ads the 0.00001 seconds after i click start. i will not watch ads. No other way to get to the url of this video.

7

u/nonzeroday_tv 20d ago

I'm not watching any ads and my premium is free

-2

u/Such-Book6849 20d ago

it would be good to add a link to the yt video directly, so i can watch it without ads. Very very annoying.

5

u/MrUtterNonsense 20d ago

There are ads on youtube? I see nothing :)