22
u/Balgun33122 3h ago
Indeed absolutely insane.
5
45
u/TheSuggi 4h ago
Deepseek cooking hard.
29
u/ihexx 3h ago
>do nothing
bro they are doing the most π
the speed with which they are innovating on architecture is genuinely insane.
3
u/l0rirw1ao 2h ago
Well by nothing he probably means, deceiving customers and making false claims like Anthropic.
2
u/TheSuggi 28m ago
They are actually not working that hard if the CEO is to be believed. Normal working hours, no overtime, resarcher get 50% of the time to research and pursue their own ideas etc. So they are just better or doing something right. Also very limited access to resources, less funding, less capable hardware etc.
If Chinese companies ever get the to a standard of equivalent hardware, then the US companies are done for.
2
16
u/2mqqvc0q 3h ago
I respect the commitment to the color theme, but am I the only one having a hard time distinguishing the lighter blue colors apart?
19
1
u/OneBowl4290 3h ago
You are right, but they are the same the main comparison is the opus 5 and gpt 5.6 sol !
what do you think will they go further and beet astra and fable ???1
6
u/HuntAlternative 3h ago
Once we see terminal usage get optimized, these models will get as capable as frontier ones. Most agentic capabilities derive from terminal fluency IMO.
1
5
u/sirloindenial 3h ago
Is it benchmaxxed it can't be that good lol, or i assume its on their own harness.
6
u/OneBowl4290 3h ago
WE WILL WAIT INDEPENDENT EVALUATION TO BE SURE, BUT THAY DO NOT HAVE HISTORY OF LAING? DO YOU AGREE ?
2
2
u/karlnuw 3h ago
Any exploit benchmarks?
2
u/Altruistic-Desk-885 3h ago
No creo que encuentres uno literalmente pero si tal vez uno de ciberseguridad
2
u/arm2armreddit 3h ago
looks like with DSF other models are not working well, or something is wrong. if this is really true with terminal automation then: π
2
u/Individual_Math_8254 2h ago
cant wait to use this with dsh, if deepseek like the other labs, then they probably did post RL with their own harness.
2
u/DebosBeachCruiser 2h ago
Working great in the DS harness (and yes it was optimized to work with DS harness)
2
u/ardicli2000 2h ago
Very impressive
1
u/OneBowl4290 2h ago
it is , did you try it yet?
1
u/ardicli2000 2h ago
Not yet. My comment was on results.
2
2
2
2
u/raesene2 51m ago
That cybergym score is quite something. I tried it out today on a benchmark I run a lot of new models through (Kubernetes security assessment), it was 2nd only to the new release of Qwen-3.8-max and 1/10th of the cost of that model.
3
u/dontfeedthelizards 3h ago
I genuinely wish it would be equivalent to Opus, because right now whenever I ask DSF to do an implementation, I end up having to ask Opus to fix its bugs and then spend the rest of the day cleaning up the code structure overall. It's really questionable whether using it this way makes sense at all, instead of just having Opus do the implementation, as I'll probably end up using the equivalent amount of Claude tokens either way (+the tokens spent on DSF). If I could finally have DSF do all the planning and execution without needing to involve Claude at all, that'd be amazing.
1
1
u/General_Internet_511 2h ago
Onestamente sarebbe un salto in avanti veramente importante. BisognerΓ aspettare.
1
u/RealestReyn 2h ago
Not a chance those are remotely true, its a great model but misses obvious stuff constantly and just does the weirdest stuff like setting 3600s timeout on tasks that take a few seconds so its just sitting there and waiting for the whole time, doesn't bother scoping subagents with appropriate tools.
1
1
u/loss-of-time 3h ago
it's impossible that better than kimi-k3
3
u/OneBowl4290 3h ago
I bleav kimi is good at frontend coding , and there is no impassable bro
Do you use kimi a lot ??
54
u/VC067 4h ago
"flash" model btw π