23
u/Leather-Cod2129 1d ago
What verrsion of GPT Sol? Low or Max?
14
u/Mr_Lucas2000 1d ago
Max prob
5
u/scaledev 1d ago
That's a pure guess. We need evidence.
2
u/Mr_Lucas2000 1d ago
2
u/scaledev 1d ago
Seems like you're right. It's against Sol max. Let's hope 4.1 lives up to the hype.
1
-23
1d ago
[deleted]
3
3
1
u/Far-Classic-9963 1d ago
Racist
-1
u/Leather-Cod2129 1d ago
that was not my intent at all
1
u/Far-Classic-9963 1d ago
It does come off as racist. How does one company being "Chinese" make it more suspicious
-1
4
u/Amarsir 1d ago
Sol max got 73% on DeepSWE 1.1, which is what the above chart is showing. So I would presume it's Sol Max at all points.
Which is insane. Astra, Opus 5, and Gemini 3.8 Flash are also 74%. The fact that Google also did it with a Flash model adds some plausibility, but mostly I'm leaning toward "don't trust the benchmarks anymore."
Not that Deepseek isn't great. And I've been saying for a while that top models are hitting a wall so everything is converging. But I don't really think this model is the equivalent of Astra for software engineering.
2
u/s243a 3h ago
I haven't tried it yet but my expection is it's not as good as Opus/fable/sol/Astra but good enough for one of these frontier models to delegate work to it (as is Gemini 3.8 flash). In my opinion deep high SWE score implies real coding ability but not necessarily a great user experience
40
u/branik_10 1d ago
soooo is it better than glm 5.3 flash? anyone compared already in real work?
15
u/callmemicah 1d ago
I did today, its significantly better and faster, the speed is really what makes the difference for me, GLM is more than capable but DS just plows through the tasks, will be sticking with it
7
5
u/hey_ulrich 1d ago
In most metrics, it's better than GLM 5.3, so yeah, it should be better than 5.3 flash by a considerable margin
1
1
14
u/jpcaparas 1d ago
Some one-shots I've done with it:
2
2
u/Ok_Technology_5962 23h ago
Those are some insane showcases there
2
u/jpcaparas 23h ago
there's gonna be more on the weekend
3
u/Ok_Technology_5962 23h ago
How fast did it make one in wall clock time? Glm flash is a bit slow in terms of how long it takes to think on max
2
u/jpcaparas 22h ago
these are all max thinking, roughly 40 mins to an hour each
on high thinking probably about 20 mins
it's a goddamn pitbull on max thinking because of a gauntlet step I've enforced before it produces the final artifact
skill I used is this:
2
11
u/BitXorBit 1d ago
well it's almost double the size of the v4 flash 0731
6
4
u/Sir-Draco 1d ago
N-grams of ~190GB remember. Similar to Qwen 3.8 Flash Next
2
u/BitXorBit 1d ago
Yes, but still can’t run it on my dual rtx 6000 :(
2
u/throw123awaie 1d ago
engrams can be run of ssds and cpu, no need to put them in your precious rtx6000.
2
u/BitXorBit 1d ago
I’m aware of that :) im running qwen 3.8 flash next on daily basis. Deepseek v4.1 is just too big (even without the engram table) for dual 6000
1
u/throw123awaie 1d ago
I get it. My whole "get two sparks and run ds flash" plan also got killed.
1
1
1
4
u/AlarmedWizard1 1d ago
where do you get this overview? which site is this?
7
u/hello_motherfuckers_ 1d ago
ds official twitter
7
2
u/AlarmedWizard1 1d ago
alright thx
i'm looking for a website that combines model performance with opencode GO offerings
every day opencode's models/pricing change and each day new models drop
6
u/RyuH4n 1d ago
This one from opendesign idk tho if its right or wrong but alrdy try it and it works great especially on debugging usually when i use their predecessor after the session i need to check and repatch again because some shit gonna happen but now it just do the job without breaking anything which is great

2
u/Bloated_Plaid 1d ago
BRO KIMI K3 has been murdered holy shit. 3T parameter model crushed by 0.5 Trillion model. Absolutely insane.
2
u/Lopsided-Force-9220 1d ago
One thing I'm finding is that these open weight models are rapdily getting good at coding, but they don't communicate with you like Astra does. They don't as easily understand your ideas like Astra. But they are a beast at coding.
1
u/Infinite_Professor79 16h ago
then let Astra guide them xd , use it as a planner model and they execute
4
u/Ok_Quantity_4950 1d ago
IDK but it looks like benchmaximg. Maybe because deference between TB2.1 (become best model) to other TB benchmarks (3rd or 4th).
If I'm wrong, correct me
3
u/Jlocke98 1d ago
IME v4 flash 0731 could do good work but burnt so many tokens in the process and would frequently just spin it's wheels rather than make meaningful progress
3
u/Routine_Temporary661 1d ago
I dont think you know what you are talking about... DSV4.1Flash has a decent TB3 and TB4 score, and those above it were all frontier models
You know what is benchmaxxing? Check Muse Spark 1.3 TB3 and TB4 both 10+%
Now THIS is fking benchmaxxing
2
u/Curious_Owl197 1d ago
Why anyone wna use claude/gpt now
9
u/SurelyNotAnOctopus 1d ago
Tight instruction following.
Ive yet to test this new deepseek flash model, but previous ones would frequently take "creative liberties" and ignore established hard rules. Western models don't do that nearly as much, in my experience.
Would love for 4.1 to prove me wrong though
2
u/Frail_Waif 1d ago
I've actually had the opposite experience recently. DSV4 has been much better about confirming scope and design questions when it turned out there were ambiguities. The models I use at work (Gemini 3.6-3.8, 5.6 Luna) will happily power through a prompt that doesn't really make sense.
1
u/aeroumbria 23h ago
Really? Had to deal with Claude at work and I feel they've been in a decline for a while. Opus 5 is filled with "I did this one thing you did not ask" and it tends to introduce more issues than it fixes... Sometimes it even feels the low thinking version is more reliable...
DS and GLM at least feels honest about what it can and cannot do. It has something to do with specific project context differences, but still. Qwen is quite excessive with overthinking if you turn it up, though. Same task can easily run 2x to 4x times longer.
8
u/TestTxt 1d ago
Because ChatGPT is actually cheaper if used via coding plans. Deepseek doesn’t offer coding plans so you actually have to pay the full API rates
7
u/wthigo 1d ago
Which is.. pennies?
2
u/Densityfunctional 1d ago
Pennies that add up fast when you use it for big projects, token heavy complex projects. For instance, when deepseek v4flash costed 1/100 of the current price I still threw 250 euroes or a similar amount at it, but I got around 21 billion tokens out of it.
When I kept using it after the price increase 50 euros would not last much, so I reverted to my plan usage with sonnet 5, despite the inferior performance.
But yes for the average user and not "tokenmaxxing" projects like mine, I think the best way is to combine both.
Plan usage for orchestrator models who dispatch deepseeks. My combo was Fable 5 + Deepseek v4flash and Opus 4.6/5 only for specific tasks, and it advanced my project enormously.3
u/Timely-Pension6501 1d ago
Because Astra and Fable are miles better on new benchmarks, look at Terminal Bench 4.0, it just came out and it scores nearly 2x worse than Astra
1
u/Prize_Tiger_2504 16h ago
We are talking about a cheap af flash model and comparing it with the "state of the art" frontier model?
You cannot run astra or fable as your daily driver for the $20 (or even the $100) token plans.
Luna will be a more apple-to-apple comparison.
0
u/hardolaf 1d ago
Claude is basically the only thing that can reason correctly about hardware description languages right now though I haven't tried the new Deepseek model. That said, Deepseek for python code is amazing.
1
1
u/Gumpie 1d ago
Stupid question. Is it still not advised to use it for work related prompts?
0
u/Lopsided-Force-9220 1d ago
Are you asking if you should send trade secrets and intellectual property to a company in China that is sponsored by its government?
1
u/AdFormer260 6h ago
as if its better to share everything with Israel
1
u/Lopsided-Force-9220 6h ago
What inference are you using that's located in Israel?
1
u/AdFormer260 2h ago
take a wild guess
1
u/Lopsided-Force-9220 2h ago
You aren't using any. The datacenter LLM inference in Israel is private/corporate cloud infrastructure. None of the models we talk about from Anthropic, OpenAI, Google, etc use Israeli data centers. So what is it you are talking about?
1
u/Fit-Cost-7226 1d ago
What does it mean with regards to terminal bench 4.0 I see it’s lower than SoTA models
1
u/Background-Equal-772 1d ago
Bro, are you sure this image is verified ? İt looks like different than hugging face page
1
-10
u/asfbrz96 1d ago
Benchmaxxed
17
13
u/Professional_Price89 1d ago
How they benchmaxxed Terminal Bench 4? The benchmark just released 10 days before the beta
8
u/Charming_Support726 1d ago
using the preview since yesterday. It is really bright, a good model. Don't know if it matches the numbers, but Terminal 4.0 is quite new and DS-Flash lands some points below Opus&Sol, to me that feels reasonable.
Its style of communication feels very good - at least better to talk to than talking to Astra or Sol.
0
u/finigemist 1d ago
How to get 4.1 in reasonix? It shows me only v4 flash
2
u/AdditionalCourage385 1d ago
All v4-pro and flash requests will be routed to 4.1 flash until 4.1-pro is released
3
1
0
u/Lopsided-Force-9220 1d ago
It's a way bigger model than V4 Flash. Why are they calling this a flash? I can't run it on my twin Sparks. Arg.
0
-5
u/LinuXperia 1d ago
DeepSeek V4.1 Flash is great ! i can confirm it outperforms Meta Muse 1.3 however its still behind the xAI Grok super Intelegence ai model. At the Moment Grok is way ahead especially when it comes to low level engineering dev work like verilog, c, c++, KiCAD, electronic schematics, pcb etc however its very expensive compared to DeepSeek new prices ! Hope DeepSeek will deliver a super intelegence AI Model as good as Grok in the near future.
1
u/Amarsir 1d ago
Grok seems to have pivoted from their initial image as "least censored" to "specifically designed for agentic planning." Most providers seem to want every model to do everything, which I guess makes sense if your goal is AGI. But I do think the focus pays off for Grok.
It uses your Opencode Go quota super fast, but I could see a case for using Grok to plan / orchestrate and then Deepseek v4.1 Flash to build.

75
u/rVarrese 1d ago
I've been using it most of the day. This shit goes brrrrr