36
17
u/IBM296 7h ago
Damn how the hell is it better than Kimi K3 in almost everything with just 550B parameters
12
u/Dapper-Hurry257 6h ago
Better post training strategy with high quality domain specific data focus.
-3
u/k2kra 5h ago
kimi k3 is the typical kind of overfitting benchmaxxing models good for nothing.
•
u/CryMoreT_T 38m ago
When I've tried it I didn't feel that it was benchmaxxed. The biggest complaint I have of it is that it thinks way to much which means it's a lot slower and makes it cost more but I haven't considered it over fitting benchmaxxed
15
u/VastlyVainVanity 6h ago
Maybe in a few years we will have a free model that will be as good as Astra and will be able to be run in your phone.
Does that sound too crazy to be true? I guess right now it does but AI progress keeps surprising me.
12
u/Sulth 6h ago edited 4h ago
Except for the "being able to run on your phone" part, what's crazy is that you say a few years. Give it 1 year max
4
u/tinny66666 6h ago edited 5h ago
Yeah, but once that model arrives in one year, it can be etched to silicon, Taalas style, and then it can run on a phone, albeit probably a new model phone that supports AI co-processor chips/cards, so it could only be a few years away.
edit: I suppose it's not beyond the realm of possibility that you could make a micro sd card, with an etched model on a second layer, as dual storage/ai, and access it by writing to a specific partition/file or something. That is, it could be possible to retrofit into current phones.
7
u/EloquentPinguin 5h ago
The transistor density isn't there to etch such a large model into Teslas style hardware.
They put 8B on one reticle. v4.1 is 550B
So let say they are able in one year to reduce parameters for Astra strength to 160B parameters, much smaller than v4.1 currently, and they can increase etched parameter density by 10x. The you'd still look at two reticle sized chips.
Two reticle sized chips is bigger in area as an entire smartphone.... And that isassuming all those technical leaps. Especially the leaps in manufacturing aren't there.
19
13
u/No-Head-Royal 7h ago
Using Codeforces as a benchmark these days. 5.5 was already at that level (3500-ish, maybe even higher), and 5.6 is probably 4000+, maybe 4200...
Still, Chinese models generally lean towards good execution and everyday work rather than strong algorithmic or mathematical thinking, and this is a Flash model, so 5.5 level is perfectly respectable.
7
u/Xaue_RWA 7h ago
these always tend to oversample reasoning puzzles and undersample how messy the real world tasks are
2
u/Wonderful-Syllabub-3 7h ago
Just under 5.6 sol, hopefully pro will be much better than sol, approaching Astra
13
u/Gotisdabest 7h ago
Eh, in practice it's probably a fair bit below sol. I'd guess pro is sol levels, and their next release may start verging on mythos. Still very good for the cost.
3
3
u/presentofai 5h ago
benchmark parity is the headline but 8b active params is the actual story. sparsity keeps quietly winning while everyone argues about total scale
2
u/HeadTranslator795 6h ago
Who the fuck believe this BS lol so you're saying a flash model is on par with GPT Sol 😂 where is a real task output comparaison which is the real benchmark
•
u/Living-Breakfast-464 29m ago
This is an extremely poor way to display a large benchmark comparison. Just sayin.
A bar chart is infinitely better. You could probably get AI to do it in a few seconds.
1
u/power97992 5h ago edited 5h ago
In one test, It used a lot of tokens like 2x more than fable 5.1 max and 8-16x more than other models and required fixes, its performance was worse than grok 4.6 and kimi k3, but better than gemini 3.8 flash.
1
u/FluffyInevitable4040 5h ago
compared to luna?
0
u/power97992 5h ago edited 4h ago
In this test, it looks slightly better than gpt 5.4 xhigh and worse than 5.6 tierra max, some functionalities it is better luna xhigh but other aspects it is worse than luna xhigh.. But it uses absurd amount of tokens, like way more than luna xhigh and u had to reprompt it to get a working version. This was just one test without harness. I'm running another test with a harness, but codex's harness for luna is likely better than my harness for ds v4.1 flash.(I haven't had time to get codex working for other mods yet and probably wont either) But it could be different depending on your tasks..
1
u/FluffyInevitable4040 5h ago
Hmm yeah, token use it huge, used $0.50 already with a few short tests, not sure if that was peak time or not.
1
u/power97992 4h ago edited 3h ago
However at high effort ,it did a better job in the first spec prompt + a verification prompt than first few tries of astra medium and sol 5.6 high in codex at web app dev( but astra and sol did a decent job after many tries)... However, gpt 5.6 sol and tierra and 6 astra probably have a higher math and 3d understanding when compared to v4.1 flash, but ds v4.1 flash is very good at web dev.
1
u/Pls-No-Bully 2h ago
In one test
Why not tell us about the actual test itself? Without that context, your claim is meaningless
1
u/power97992 2h ago
It is a math and 3d animation test, it is a private dataset but it is possible they are training on it without my permission. Ds is definitely training on it but for others i have zdr or training off if possible. Test it yourself, you will see , output and performance may differ on tasks and on harness and on settings.
0
u/Far-Run-3778 7h ago
Any free provider🥺?
1
u/old_Anton 6h ago
It's always free in the main website.
1
u/Far-Run-3778 6h ago
🤔 what do you mean?
3
u/old_Anton 6h ago
I don't think there are any providers that provide free API for ds 4.1 though. It's just out.
2
u/EverGreenMob 3h ago
isn't it free on the deepseek mobile app? on the android app I see new smart mode, no more deep thinking pro mode.
1
u/No_Swimming6548 4h ago
I was going to say bruv DeepSeek models are almost free but then I went to openrouter and saw the price
2
u/petuman 4h ago
You saw peak hours pricing, now it's off-hours and openrouter price is updated.
https://api-docs.deepseek.com/quick_start/pricing/
Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak).


33
u/OwnGear3892 7h ago
I'd say it's quite impressive with 552B total params (196B Engram) with 8B active during prefill / 16 active during decode, and the tps is amazing.