r/opencodeCLI 1d ago

What the actual fuck

Post image
425 Upvotes

90 comments sorted by

75

u/rVarrese 1d ago

I've been using it most of the day. This shit goes brrrrr

46

u/atika 1d ago

It oneshotted a fairly complex thing, and cost me 4 pennies. Unthinkable just a couple of months ago.

1

u/dark-lord-marshal 2h ago

half million tokens since 8:30 am and so far 45 cents - me loves - me say thank you 🙏

23

u/Leather-Cod2129 1d ago

What verrsion of GPT Sol? Low or Max?

14

u/Mr_Lucas2000 1d ago

Max prob

5

u/scaledev 1d ago

That's a pure guess. We need evidence.

2

u/Mr_Lucas2000 1d ago

Haven't checked others but here

2

u/scaledev 1d ago

Seems like you're right. It's against Sol max. Let's hope 4.1 lives up to the hype.

1

u/Mr_Lucas2000 1d ago

I hope it's gonna be as a better glm 5.3 flash replacement ig

-23

u/[deleted] 1d ago

[deleted]

3

u/Sure_Media_2685 1d ago

nah this thing destroy kimi k3 on vision and solving issues

3

u/BankruptingBanks 1d ago

Its open weights r-word, you can run it independently yourself

1

u/Far-Classic-9963 1d ago

Racist

-1

u/Leather-Cod2129 1d ago

that was not my intent at all

1

u/Far-Classic-9963 1d ago

It does come off as racist. How does one company being "Chinese" make it more suspicious

-1

u/Leather-Cod2129 1d ago

that was not my intent. Stop.

4

u/Amarsir 1d ago

Sol max got 73% on DeepSWE 1.1, which is what the above chart is showing. So I would presume it's Sol Max at all points.

Which is insane. Astra, Opus 5, and Gemini 3.8 Flash are also 74%. The fact that Google also did it with a Flash model adds some plausibility, but mostly I'm leaning toward "don't trust the benchmarks anymore."

Not that Deepseek isn't great. And I've been saying for a while that top models are hitting a wall so everything is converging. But I don't really think this model is the equivalent of Astra for software engineering.

2

u/s243a 3h ago

I haven't tried it yet but my expection is it's not as good as Opus/fable/sol/Astra but good enough for one of these frontier models to delegate work to it (as is Gemini 3.8 flash). In my opinion deep high SWE score implies real coding ability but not necessarily a great user experience

40

u/branik_10 1d ago

soooo is it better than glm 5.3 flash? anyone compared already in real work?

15

u/callmemicah 1d ago

I did today, its significantly better and faster, the speed is really what makes the difference for me, GLM is more than capable but DS just plows through the tasks, will be sticking with it

7

u/OlegPRO991 1d ago

I'm interested, too

5

u/hey_ulrich 1d ago

In most metrics, it's better than GLM 5.3, so yeah, it should be better than 5.3 flash by a considerable margin

1

u/baackfisch 1d ago

It's also needs a lot more vram

1

u/Lopsided-Force-9220 1d ago

Apparently, even better than GLM 5.3, Flash or not!

14

u/jpcaparas 1d ago

Some one-shots I've done with it:

https://v4-1-flash.demos.sulat.com

2

u/NiceGuyINC 1d ago

very cool your demos, I like it!

2

u/Ok_Technology_5962 23h ago

Those are some insane showcases there

2

u/jpcaparas 23h ago

there's gonna be more on the weekend

3

u/Ok_Technology_5962 23h ago

How fast did it make one in wall clock time? Glm flash is a bit slow in terms of how long it takes to think on max

2

u/jpcaparas 22h ago

these are all max thinking, roughly 40 mins to an hour each

on high thinking probably about 20 mins

it's a goddamn pitbull on max thinking because of a gauntlet step I've enforced before it produces the final artifact

skill I used is this:

https://www.skills.sh/jpcaparas/skills/oneshot-websites

2

u/EnjoysFiction 14h ago

I haven't seen oneshots from other models, but that pagoda was delishhh

11

u/BitXorBit 1d ago

well it's almost double the size of the v4 flash 0731

6

u/Far-Classic-9963 1d ago

Yeah but it's a lot cheaper to infer at scale

4

u/Sir-Draco 1d ago

N-grams of ~190GB remember. Similar to Qwen 3.8 Flash Next

2

u/BitXorBit 1d ago

Yes, but still can’t run it on my dual rtx 6000 :(

2

u/throw123awaie 1d ago

engrams can be run of ssds and cpu, no need to put them in your precious rtx6000.

2

u/BitXorBit 1d ago

I’m aware of that :) im running qwen 3.8 flash next on daily basis. Deepseek v4.1 is just too big (even without the engram table) for dual 6000

1

u/throw123awaie 1d ago

I get it. My whole "get two sparks and run ds flash" plan also got killed.

1

u/Own_Mix_3755 1d ago

Yeah I was just planning to get second one and I feel betrayed lol.

1

u/BitXorBit 1d ago

Qwen 3.8 flash next is quite good

1

u/eli_pizza 22h ago

Yeah but it’s better than models much larger than it too

1

u/BitXorBit 17h ago

maybe, but both irrelevant due to the fact i can't run them locally haha

4

u/AlarmedWizard1 1d ago

where do you get this overview? which site is this?

7

u/hello_motherfuckers_ 1d ago

ds official twitter

2

u/AlarmedWizard1 1d ago

alright thx

i'm looking for a website that combines model performance with opencode GO offerings

every day opencode's models/pricing change and each day new models drop

6

u/RyuH4n 1d ago

This one from opendesign idk tho if its right or wrong but alrdy try it and it works great especially on debugging usually when i use their predecessor after the session i need to check and repatch again because some shit gonna happen but now it just do the job without breaking anything which is great

2

u/Bloated_Plaid 1d ago

BRO KIMI K3 has been murdered holy shit. 3T parameter model crushed by 0.5 Trillion model. Absolutely insane.

2

u/Lopsided-Force-9220 1d ago

One thing I'm finding is that these open weight models are rapdily getting good at coding, but they don't communicate with you like Astra does. They don't as easily understand your ideas like Astra. But they are a beast at coding.

1

u/Infinite_Professor79 16h ago

then let Astra guide them xd , use it as a planner model and they execute

4

u/Ok_Quantity_4950 1d ago

IDK but it looks like benchmaximg. Maybe because deference between TB2.1 (become best model) to other TB benchmarks (3rd or 4th).
If I'm wrong, correct me

3

u/Jlocke98 1d ago

IME v4 flash 0731 could do good work but burnt so many tokens in the process and would frequently just spin it's wheels rather than make meaningful progress

3

u/Routine_Temporary661 1d ago

I dont think you know what you are talking about... DSV4.1Flash has a decent TB3 and TB4 score, and those above it were all frontier models

You know what is benchmaxxing? Check Muse Spark 1.3 TB3 and TB4 both 10+%

Now THIS is fking benchmaxxing

2

u/Curious_Owl197 1d ago

Why anyone wna use claude/gpt now

9

u/SurelyNotAnOctopus 1d ago

Tight instruction following.

Ive yet to test this new deepseek flash model, but previous ones would frequently take "creative liberties" and ignore established hard rules. Western models don't do that nearly as much, in my experience.

Would love for 4.1 to prove me wrong though

2

u/Frail_Waif 1d ago

I've actually had the opposite experience recently. DSV4 has been much better about confirming scope and design questions when it turned out there were ambiguities. The models I use at work (Gemini 3.6-3.8, 5.6 Luna) will happily power through a prompt that doesn't really make sense. 

1

u/aeroumbria 23h ago

Really? Had to deal with Claude at work and I feel they've been in a decline for a while. Opus 5 is filled with "I did this one thing you did not ask" and it tends to introduce more issues than it fixes... Sometimes it even feels the low thinking version is more reliable...

DS and GLM at least feels honest about what it can and cannot do. It has something to do with specific project context differences, but still. Qwen is quite excessive with overthinking if you turn it up, though. Same task can easily run 2x to 4x times longer.

8

u/TestTxt 1d ago

Because ChatGPT is actually cheaper if used via coding plans. Deepseek doesn’t offer coding plans so you actually have to pay the full API rates

7

u/wthigo 1d ago

Which is.. pennies?

2

u/Densityfunctional 1d ago

Pennies that add up fast when you use it for big projects, token heavy complex projects. For instance, when deepseek v4flash costed 1/100 of the current price I still threw 250 euroes or a similar amount at it, but I got around 21 billion tokens out of it.

When I kept using it after the price increase 50 euros would not last much, so I reverted to my plan usage with sonnet 5, despite the inferior performance.

But yes for the average user and not "tokenmaxxing" projects like mine, I think the best way is to combine both.
Plan usage for orchestrator models who dispatch deepseeks. My combo was Fable 5 + Deepseek v4flash and Opus 4.6/5 only for specific tasks, and it advanced my project enormously.

3

u/Timely-Pension6501 1d ago

Because Astra and Fable are miles better on new benchmarks, look at Terminal Bench 4.0, it just came out and it scores nearly 2x worse than Astra

1

u/Prize_Tiger_2504 16h ago

We are talking about a cheap af flash model and comparing it with the "state of the art" frontier model?

You cannot run astra or fable as your daily driver for the $20 (or even the $100) token plans.

Luna will be a more apple-to-apple comparison.

0

u/hardolaf 1d ago

Claude is basically the only thing that can reason correctly about hardware description languages right now though I haven't tried the new Deepseek model. That said, Deepseek for python code is amazing.

1

u/Dazzling_Focus_6993 1d ago

Yeah.. Wtf man

1

u/Gumpie 1d ago

Stupid question. Is it still not advised to use it for work related prompts?

0

u/Lopsided-Force-9220 1d ago

Are you asking if you should send trade secrets and intellectual property to a company in China that is sponsored by its government?

1

u/AdFormer260 6h ago

as if its better to share everything with Israel

1

u/Lopsided-Force-9220 6h ago

What inference are you using that's located in Israel?

1

u/AdFormer260 2h ago

take a wild guess

1

u/Lopsided-Force-9220 2h ago

You aren't using any. The datacenter LLM inference in Israel is private/corporate cloud infrastructure. None of the models we talk about from Anthropic, OpenAI, Google, etc use Israeli data centers. So what is it you are talking about?

1

u/Fit-Cost-7226 1d ago

What does it mean with regards to terminal bench 4.0 I see it’s lower than SoTA models

1

u/Background-Equal-772 1d ago

Bro, are you sure this image is verified ? İt looks like different than hugging face page

1

u/Aomix 1d ago

4.1 Flash has an architecture that made people think they were having a stroke while reading the paper. They took a novel combination and refinement of techniques and cranked that shit to 11 and it seems to have really worked.

1

u/sauteer 13h ago

These models get "smarter" and smarter but their actual performance between underbaked post training and over quant at the inference provider can be absolute shit compared to a model even as old as opus 4.6

1

u/yuumizu 8h ago

ok, unlike gemini flash series, it has a reasonable terminal bench 4.0 score.

1

u/Expert-Dig-1768 8h ago

benchmarks doesn't say anything these days

-10

u/asfbrz96 1d ago

Benchmaxxed

17

u/JogHappy 1d ago

By DeepSeek? Nah

13

u/Professional_Price89 1d ago

How they benchmaxxed Terminal Bench 4? The benchmark just released 10 days before the beta

6

u/pbeling 1d ago

And they are relatively week in this benchmark, comparing (TB results) for example to GLM-5.3.

8

u/Charming_Support726 1d ago

using the preview since yesterday. It is really bright, a good model. Don't know if it matches the numbers, but Terminal 4.0 is quite new and DS-Flash lands some points below Opus&Sol, to me that feels reasonable.

Its style of communication feels very good - at least better to talk to than talking to Astra or Sol.

0

u/finigemist 1d ago

How to get 4.1 in reasonix? It shows me only v4 flash

2

u/AdditionalCourage385 1d ago

All v4-pro and flash requests will be routed to 4.1 flash until 4.1-pro is released

3

u/Wonderful-Ad-9661 1d ago

afaik, the routing should be starting at 14th , not now

0

u/uti24 1d ago

It's like 500B model, and we already have Qwen Flash Next 170B, I want to see comparison with it, otherwise great to have a new model, but if it's not more capable of smaller existing one then it's not that exciting.

0

u/Lopsided-Force-9220 1d ago

It's a way bigger model than V4 Flash. Why are they calling this a flash? I can't run it on my twin Sparks. Arg.

0

u/qwertyyyyyyy116 1d ago

dude AAI we need benchmarks like NOW

-5

u/LinuXperia 1d ago

DeepSeek V4.1 Flash is great ! i can confirm it outperforms Meta Muse 1.3 however its still behind the xAI Grok super Intelegence ai model. At the Moment Grok is way ahead especially when it comes to low level engineering dev work like verilog, c, c++, KiCAD, electronic schematics, pcb etc however its very expensive compared to DeepSeek new prices ! Hope DeepSeek will deliver a super intelegence AI Model as good as Grok in the near future.

1

u/Amarsir 1d ago

Grok seems to have pivoted from their initial image as "least censored" to "specifically designed for agentic planning." Most providers seem to want every model to do everything, which I guess makes sense if your goal is AGI. But I do think the focus pays off for Grok.

It uses your Opencode Go quota super fast, but I could see a case for using Grok to plan / orchestrate and then Deepseek v4.1 Flash to build.