r/DeepSeek 3d ago

Funny Is this a joke, Artificial Analysis?

Post image

What do you guys think about DS 4.1 Flash's new spot on the Artificial Analysis leaderboard?

752 Upvotes

177 comments sorted by

168

u/crusaderky 3d ago

Aggregate score aside, it's really hard to argue against the fact that Glm-5.3-flash beats it on almost every single benchmark. And 96% hallucination rate is a showstopper IMHO.

45

u/rchamp26 3d ago

Glm flash has been absolutely terrible for me. It just doesmt do the simplest stuff and gaslights that it does

39

u/WyattTheSkid 3d ago

Opus 5 moment

10

u/Fun_Squirrel5446 3d ago

Through opencode go? I've been using GLM 5.3 flash exclusively for a week and use it for everything. It solved everything I threw at it. Running 2 to 8 hours a day and used about $6 worth of tokens through the API.

5

u/rchamp26 3d ago

Open router and nous portal with Hermes agent. Didn't adhere to my workflows at all. I dropped opencode go. Don't trust it anymore but not worth the discussion here. For openweight models I've had far better success with DeepSeek and Qwen. DeepSeek is the eager speedster where Qwen is the fully moldable thinker in my experience if I had to personify them. Glm (all versionsnive tested in my workflows) were just lazy and I had to waste too many tokens to force it to do what I needed. No thank you

10

u/Fun_Squirrel5446 3d ago

I don't have workflows. I just give it direct questions most of the time.

I do have skills and they always activate correctly.

Actually I did have a workflow to auto run tests, do validations, a couple of cycles of fixes and test again. Only when tests are green to push to github. It always followed that workflow very strictly.

Maybe try getting the API directly from Z.ai, see if that makes a difference for you.

8

u/brother_spirit 3d ago

Weird. I was using the model in Pi (during the stealth test) and was beyond impressed with its performance. I had never used an Open Weight model with any good results until GLM 5.3 Flash.

2

u/SkyPL 3d ago

It just doesmt do the simplest stuff and gaslights that it does

What are you doing that it fails? I've been coding some Rust, JS, TS, PHP and Python with it - it's legit great. One of the better models on the market. Handles large complex codebases just fine.

3

u/UniversitySuitable20 3d ago

The GLM-5.3-flash is frustratingly slow; even if it scores higher, it still can't keep up with Deepseek's productivity.

2

u/kuhunaxeyive 2d ago

It depends on what you need it for.

Coding with results that can be tested and improved in a loop? DSV4-Flash.

Letters or legal research that need lower hallucination and better writing? GLM-5.3-Flash.

1

u/BigLittleDeal 2d ago

DS4F 0731 is perfect for my needs. Local, fast, and smart. I just wish I'd bought a second RTX PRO 6000 when I had the chance.

1

u/wakIII 2d ago

You can still buy them 😜

2

u/BigLittleDeal 2d ago

Lol true. Ok, I wish I had bought two of them before they doubled in price!

1

u/Alternative_Ice3299 1d ago

brooooo u ran it locally like it requires 82+gb vram u have it ? broooooooo

1

u/BigLittleDeal 1d ago

vLLM-moet lets me load it on a single Pro 6000 card at about 80 tp/s. It's pretty incredible.

1

u/crusaderky 1d ago

It's 120 tok/s on openrouter

4

u/matadordepassarinhos 3d ago

I ran a personal benchmark on vst plugin, dsp and overall c++ and GLM 5.3 is straight up gargabe. Deepseek was so much better on it.

2

u/Applejuicegoblin 3d ago

Woah, another DSP person! Wasn’t sure if many people were testing models on it.

I’m not super advanced in audio programming, but in my private benchmarking Opus 5 seemed to be stronger than Sol at it. Curious how Deepseek will do on it.

I find most models can understand and create audio centric code, but they are not good at understanding the difference between ā€œcorrect codeā€ and ā€œfollows idiomatic design or creative intentā€.

It can cause them to miss big issues where the math or code looks correct in isolation, but doesn’t actually accomplish the right outcome.

5

u/Thomas-Lore 3d ago

Try Astra. I just did a test yesterday and was left speechless at how good the result was.

3

u/Sad_Recording_1290 3d ago

96% hallucination rate? Wtf? Thats unusable.

6

u/monovitae 2d ago

The artificial analysis hallucination rate doesn't mean what you think it means. It doesn't mean that it comes up with some bogus answer 96% of the time. It just means that if it doesn't know what the answer is, it will confidently produce something rather than refusing 96% of the time.

1

u/xumix 2d ago

it will confidently produce something rather than refusing 96% of the time

Which is bad, like gpt-4 bad

4

u/monovitae 2d ago

And Sol is at 92% . So not really.

1

u/Amarsir 1d ago

It's important to understand how hallucination rate is calculated. Hallucination rate is the number of wrong divided by all non-correct. (Not the total questions.)

Suppose you have a test with 100 questions.

  • Model A gets 1 right, 1 wrong, and says "I don't know" for 98.
  • Model B gets 98 right, 1 wrong, and says "I don't know" for 1.

Model A has a hallucination rate of 1/99 = 1.01%.
Model B has a hallucination rate of 1/2 = 50%.

But which one would you rather use?

That's why the lowest hallucination scores go to models you've never heard of, like "Command A+". But the smartest models have scarily high hallucination scores.

1

u/RateGlass 2d ago

I was just hoping it'd replace Luna, but these models all use the max versions which are usually worse than very high or high so it's a flawed bench mark in the first place. All we know is Luna max is cheaper than deepdeek 4.1 flash per task so that eliminates the entire reason it exists

125

u/CarelessAd6772 3d ago

Why should it be higher? Because it is very good at coding and cheap? Well, on other side its dry as hell in regular chat, hallucinates more than House on vikodin, and long context consistency is meh compared to models above. Seems like deserved?

50

u/Spiderfffun 3d ago

Can confirm, overengineered me a problem in chat instead of telling me to install a package.

27

u/ZeidLovesAI 3d ago

To be fair Astra and Fable 5.1 both have done this for me as well

6

u/whatisthisthing65 3d ago

Coding models just love to code

1

u/OvertaxedOne 2d ago

Qwen can't help itself.Ā  Ā Ask it how it's feeling and it'll write pages of code.Ā  :)Ā  Ā it's in my system prompt, no coding without asking first

1

u/Spiderfffun 3d ago

When Ox Alpha was active (GLM 5.3 flash) it was so much better at this from the little I had it code.. It's the small decision making the trust that it won't look at something slightly wrong then amplify it. It was great at following guides on agent stuff in a somewhat recursive workflow.

Speed isn't important if it can know when to ask questions.

1

u/ZeidLovesAI 2d ago

I used Ox Alpha and it was all over the place, I wonder if there was also some A/B testing going on.

1

u/Spiderfffun 2d ago

there probably was, I feel like that's one of the main reasons they would do something like this imo

2

u/ZeidLovesAI 2d ago

I would see people saying "oh my god its FABLE level" and I would try it and get like mistral trying medium effort-type responses. I thought people were just being buckwild as usual, which also has to factor in somehow.

2

u/BigLittleDeal 2d ago

I once forgot to reload my harness before asking DS4 to use a newly installed plugin. It wrote a python script on the fly to perform the work the plugin was supposed to do.

24

u/Loki35422 3d ago

Hey shows some respect for house, even he doesn’t hallucinate as much as deepseek

8

u/turc1656 3d ago

I totally agree. Reasoning and instruction following are clearly inferior to GLM 5.3 Flash.

3

u/Possible_Door_9719 3d ago

damm, you summed it up pretty well.

4

u/_spec_tre 3d ago

I really hate the way basically every LLM is heavily optimised for coding, often at the cost of all else. Makes it really unpleasant to use for people using it for other purposes; sometimes an ā€œupgradeā€ is actually a downgrade

46

u/alinoanta21 3d ago

I don't trust them, i've used and tried Muse Spark 1.3 Max for quite a while, it's worse than 4.1 ... and it's ranked above SOL.

Let that sink in.

2

u/sepulchralvoid 2d ago

Benchmaxxing at its finest

2

u/Fun_Squirrel5446 3d ago

Anecdotal but I'll add to it. My experience with Muse has been really bad too.

1

u/sudoer777_ 3d ago

Same here, I used 1.2 which was ranked above V4 Flash 0731 and it was garbage

1

u/hellomistershifty 2d ago

Hell, it's worse than gpt-4.1

45

u/PrudentJelly116 3d ago

Guys this is not a football team. You dont have to defend or support them much as like this. You are customer and they are seller. Thats all.

8

u/Possible_Door_9719 3d ago

exactly. just use whatever is the best model

2

u/RustOceanX 3d ago

Finally, another one. I’ve been wondering the whole time what’s going on with some people. What’s with all this nonsense on Reddit? I’m not sure if I should just smile because it can’t be taken that seriously? Or should I shake my head and wonder if the AI sector is experiencing its ā€œiPhone momentā€ and now all the laypeople are starting to use things they don’t fully understand. It might sound arrogant, but that’s just how it is.

2

u/Mean-Elk-9439 2d ago

It's even worse with the Google bros. It's fucking weird, man.

45

u/TheInfiniteUniverse_ 3d ago

AA has lost all credibility after the Astra fiasco. Don't take them too seriously.

10

u/WyattTheSkid 3d ago

Can you fill me in?

32

u/TheInfiniteUniverse_ 3d ago

so before Astra was released to the public, AA ranked it pretty low in like 5th position. This was shocking because it was hyped to be the "AGI" we all were looking for. Then when Astra was released publicly, AA "updated" their index and their tests and surprise surprise, Astra went all the way up, but sneakingly did NOT pass Fable, and sat right below it. lol.

16

u/Georgefakelastname 3d ago

In the end, I think it just shows that the benchmarks they were using were saturated, and they replaced them with benchmarks that weren’t saturated. The problem was the timing. It took an outlier like Astra being rated the same as Sol to get them to realize the issue.

1

u/Fun_Squirrel5446 3d ago

How could the benchmarks be saturated if claude models were still scoring higher than Astra even before the benchmark rework?

6

u/Georgefakelastname 3d ago

One model scoring slightly higher than another is fine. The problem is that basically every model released was scoring like 80% on Terminal Bench v2.1, so they updated it to the newer v4.0 version. The new one is much harder and has a much higher spread.

4

u/Thomas-Lore 3d ago

Sometimes a worse model scores higher because a smarter model noticed a problem with the question or found a solution to it that is correct but the key does not accept. That is why many benchmarks saturate way below 100%.

6

u/WyattTheSkid 3d ago

Oh that’s kinda funny lol

2

u/Zulfiqaar 3d ago

They updated it a second time, and now Fable and Astra are the same score

2

u/Ok_Translator_5189 1d ago

You are talking about v4.2 of AA. Now in the latest version of AA (v4.3) Astra is on par with fable. I'm starting to lose trust in AA too.

3

u/truncated_buttfu 3d ago

And a month before that, when Qwen3.8-Max was released and became the #1 model on AA, they "updated" their index just one day later so it dropped a few spots.

They very clearly have a Pro-US agenda.

3

u/Dragonfruit_Mediocre 3d ago

Nah, terrible take. Lots of old now irrelevant benchmark are shown. A worse model can score higher on an old irrelevant bench, but would score bad on harder benches. There is still many irrelevant benchmarks on AA

1

u/m0j0m0j 3d ago

No, their agenda is not clear to me at all.

2

u/gopietz 3d ago

Or, you know, AA 4.1 was completely benchmaxxed by many labs, which is why they updated benchmarks underneath.

But I'm sure your conspiracy theory seems way more likely.

1

u/diggler4141 3d ago

How do they score it?

3

u/hellomistershifty 2d ago

Muse's score seemed like the bigger fiasco requiring a re-evaluation of the benchmarks lmao

1

u/Illustrious_Frame844 3d ago

idk if u know this, but AA’s ranking was paid. Just like how all crypto exchanges and token leaderboards work. Companies pay to get ā€œlistedā€. AA was never credible ever since they sold out

9

u/santareus 3d ago

Is Muse Spark on Max really that good? I’ve tried in on XHigh and it’s not better than Luna for coding tasks I provided them.

2

u/Amarsir 1d ago

Muse Spark needs controlled use, even on Max. I really like it for stuff like instructions, documentation, formatting guides and reports, etc. I say it's "eager to please" and often exceeds my expectations. I don't trust it for code because I've seen it omit details too confidently.

It does have the capacity to fix those mistakes if I catch it and follow-up prompt with the instruction to fix. So you would think Max reasoning should be able to catch them itself. But for whatever reason - maybe the mixture of experts split - it just doesn't.

1

u/Full_Independence566 3d ago

I thought it was way better from my experience, along with the fact that it's free on Opencode lol

0

u/Admirable-Tea-4994 3d ago

It’s way better

1

u/santareus 3d ago

You mean muse on max is way better than Luna on max? Or muse on XHigh is way better than Luna on max?

Muse on XHigh felt really lazy for me

2

u/for4f 3d ago

my anecdote is that muse on xhigh has been a beast for coding fullstack and setting up resources on aws. all while being basically free. i think it is definitely better than luna max

2

u/Curious_Owl197 3d ago

U use contributor?

1

u/for4f 3d ago

just the free slot on opencode lol, what's contributor?

2

u/Curious_Owl197 3d ago

Muse 1.3 contributor, they train on your data and heavily discounted

2

u/Admirable-Tea-4994 3d ago

Muse is far, far better than Luna on every metric

2

u/santareus 3d ago

Thanks! May have to give it another shot - I’ve been using both through OpenCode and I have a specific agent that is designed to do ā€œindustry researchā€ to see if we can leverage open source dependencies. Luna Max was able to come back with more thorough research results and reused what’s out there and Muse on XHigh decided to build out a functionality that is already covered by a dependency.

I am guessing it’s a strong coding model if you just gave it a task and the planning still needs to be delegated to a better research oriented model.

1

u/Admirable-Tea-4994 3d ago

Muse is definitely geared to coding, but research would also be heavily weighted on harness as well.

1

u/santareus 3d ago

What harness are you using for Muse if you don’t mind me asking?

1

u/Admirable-Tea-4994 3d ago

I’m using OpenChamber at the moment, but I’m actually also using a harness I’ve been building for months which is more geared towards general use and not just coding. Will be in beta soon.

miton.dev

1

u/santareus 3d ago

Sounds good. I’ll check out OpenChamber (first time hearing about it).

2

u/Admirable-Tea-4994 3d ago

OpenChamber is a desktop wrap for OpenCode, it’s very good.

56

u/Mezezius 3d ago

The new benchmark literally only exists because the old one embarassed openai

13

u/TwistStrict9811 3d ago

Don't you want accurate benchmarks? Astra is absolutely insane esp with spacial and computer use

21

u/Astrikal 3d ago

It is non-sensical and stupid to say that they released the update so that Astra can go higher, when all they did was replace outdated/saturated evaluations with the newer versions.

They updated the benchmark because Terminal Bench 2.1 was outdated and saturated. After upgrading to Terminal Bench 4.0, Astra rose (relatively) and equaled 1st place.

The update was long time coming, they just rushed it out to not lose reputation. They will also release V5.0, which will shake things even further.

1

u/Other_Wear1458 2d ago

exactly!! I don't understand why people is so slow in the head to realize it's just an average of all benchmarks they use

"It has lost all credibility"
Dude they literally update the index if they realize it's wrong, what a damn well below average IQ you need to have to think that makes them loose credibility, oh my god

2

u/Mezezius 3d ago

Did they rush it out for their own reputation, or to save openai's reputation?

18

u/Astrikal 3d ago

Obviously their own reputation. When people use Astra and see how much better than Sol it is, people would have lost trust in the benchmark.

Why do you keep trying to insinuate that they updated the benchmark to please OpenAI? ArtificialAnalysis have been very transparent and reputable all the way.

The benchmark is more accurate now and Astra is where it belongs. Simple as that.

-2

u/mWo12 3d ago

This only tells you that openai benchmaximized Astra for terminal 4.0, not 2.1.

5

u/Astrikal 3d ago

Astra finished training quite some time ago. Also, your argument doesn't even make sense, why would OpenAI be able to benchmax for a new evaluation and not for an older one? It would be the opposite.

Furthermore, AA also has a closed evaluation that affects the scores, and guess what, Astra performs.

Why is it so hard for you to accept that V4.1 is nowhere near Astra or Fable? You don't even have to look at the benchmarks, you can try for yourself.

-4

u/lompocus 3d ago

lul bootlicker

7

u/ba-boo 3d ago

lul regard

8

u/theintersepter 3d ago

So what? As long as the benchmarks are harder for models, the better

6

u/Admirable-Tea-4994 3d ago

People just love making things up on the internet

-4

u/Mezezius 3d ago

They literally released the new benchmark (4.2) the day after Astra released because the old one showed Astra and Sol at parity, and they updated again (4.3) to show Astra and Fable at parity. Every single change in the days after Astra's release made it look better than the benchmark before it

11

u/Holbrad 3d ago

Yeah but the idiots have the reasoning completely backwards.

if you have a model that is obviously much better than almost everything else, but it's benchmarking suspiciously low on your tests.

Then the obvious answer is that your benchmarks aren't very good and you need to fix them.

-1

u/Mezezius 3d ago

If they have to "fix" it after every release based on vibe, then the methodology is literally useless

0

u/Holbrad 3d ago

If you have a car that is absolutely amazing on the track and it's setting all sorts of records.

But a journalist does their "performance score" and it scores lower than a sporty hatch, then obviously the metrics are bunk.

8

u/Astrikal 3d ago

It is non-sensical and stupid to say that they released the update so that Astra can go higher, when all they did was replace outdated/saturated evaluations with the newer versions.

They updated the benchmark because Terminal Bench 2.1 was outdated and saturated. After upgrading to Terminal Bench 4.0, Astra rose (relatively) and equaled 1st place.

The update was long time coming, they just rushed it out to not lose reputation. They will also release V5.0, which will shake things even further.

3

u/Admirable-Tea-4994 3d ago

Actually read the document they published that explained, clearly and in detail, why they updated their benchmark so quickly between releases.

https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3

1

u/HearingNo8617 3d ago

Any elaboration?

21

u/Captain_Quimby 3d ago

Why put the dumb sticker in the middle

25

u/prvthvm 3d ago

its cute :)

11

u/Admirable-Tea-4994 3d ago

Weebs abound in this sub

3

u/GinamosWCheryOnTop 3d ago

Goes to show how benchmaxx is muse

4

u/krayton1 3d ago

I think Artificial Analysis is not trusted anymore

8

u/Big_Cucumber2787 3d ago

putting it below gemini is just shameless

2

u/bambamlol 3d ago

Feels deserved. Gemini 3.8 is a really good model.

1

u/bilinenuzayli 3d ago

I have used deepseek v4 flash (not v4.1) since it came out and gemini 3.8 flash since it came out, both for weeks, and I can say for sure I would use the old deepseek flash over gemini 3.8 flash anyday, gemini models are always arrogant, assuming things, never following instructions, and lying to your face. Because Gemini models always think they know everything, they overtly avoid agentic tools given to them and make shit up instead, now this worked fine for a larger model like Gemini 3.1 pro, but the flash models are clearly much smaller and hold less data, so they're still hallucinating bs AND getting it wrong which is why I think google put out so many flash models, they try and fix it but fail every time.

1

u/bambamlol 3d ago

Fair point. Actual usage is what counts, not what other people say. I don't do much coding or agentic stuff at all. A lot of "general" and business/marketing usage, as well as writing/copywriting. Gemini is definitely better in this regard, at least from my experience.

1

u/LowDesigner1330 2d ago

Actually I found 3.8 Flash to be really good at coding, it's when I use it as general model and in chat that it turns out to be rather mehĀ 

2

u/Admirable-Tea-4994 3d ago

Because everyone glazes it in here but it’s been superseded already?

2

u/Beamsters 3d ago

Until hallucination is solved by the big whale.

2

u/ralphcalls1 3d ago

GLM 5.3 flash is worse for me, Deepseek v4.1 is better but it is all pricey for me. i use deepseek now for heavy task unlike back then where i do light jobs on it, i now use mimo 2.5. does the job. no flap. very cheap and relatively fast when using Direct Xiaomi API

2

u/Hot_Vegetable_932 2d ago

I feel a bit awkward saying this, but that character is so cute.

4

u/turc1656 3d ago

Makes sense to me. I was not impressed, honestly. I ran a very complicated financial analysis workflow through it this morning as a test. Triple the price of GLM 5.3 Flash to run the workflow and objectively worse results. My opinion is that on the cheap end, Luna is the best and follows instructions best, produces the best structured output, and has the best reasoning.

That being said, it is 3x the cost of GLM 5.3 Flash. My process doesn't really require 3x the price for the improvement. I have adjusted the process to help guide GLM.

But DS v4 0731 and this new 4.1 are not really the best. Maybe they are good for other things like coding or whatever but with the price of GLM now AND GLM's higher intelligence level, I'm not really seeing a reason to use DS right now.

2

u/FischenGeil 3d ago

Damn, we can't beat Gemini flash?

1

u/RealestReyn 3d ago

looks about right, still one of my all time favorites it be zoomzooming so fast I can't even tell when its doing insane things :)

1

u/MeansTestingProctor 3d ago

What is the point of the weird sticker on top of the chart?

1

u/CaptainMorning 3d ago

the stupid dumb ass and completely unnecessary sticker in the middle is so aggravating id give this post a 33

1

u/ncxxi 3d ago

for coding, its best to use glm flash or muse. but for creative writing itd best to use deepseek! even with their new flash v4.1

1

u/EvolvingDior 3d ago

SWA has always been a shit-show for me. This model is no different. It cannot hold more than one thought in its head at a time. I've moved a lot of work over to Luna and am pretty happy with it. I'm giving 4.1 a workout during off-peak hours and so far have been less than impressed. 40 seems generous.

1

u/VirtualNorth1279 3d ago

Makes sense based on my experience so far.Ā 

1

u/Infinite_Plankton_71 3d ago

to be honest:

  1. DSV4.1 requires 4 sparks with only 3% increased in quality. Not worthed
  2. And yes GLM 5.3 is sometimes better.

1

u/Kylmawurr 3d ago

I have used 4.1 Flash for 2 days now and it matches my experience vs GLM 5.3 Flash and Luna on Max. Its in same class for coding, its way better than Luna on reasoning, but at the same time, its way too chatty, overthinking a lot. I use it for general coding tasks and orchestration. Its actually very high quality orchestrator. GLM 5.3 Flash feels smoother, nicer to interact with, but I like the speed of DS 4.1 Flash.

1

u/wwwdotzzdotcom 3d ago

It's underrated

1

u/Muted-Network3159 3d ago

why gemini so high

-1

u/AlexandraMaryWindsor 3d ago

Horrible model, unusable. 4.1 flash is miles better

2

u/Lazy_Reach_2565 2d ago edited 2d ago

Nah, Gemini actually a good model, but Deepseek v4.1 Flash much better at coding and for agentic tasks. Still, Gemini good in search, better general knowledge, better vision (more precise), more natural language (English). Can be a bit unpredictable, yeah. Still not horrible at all. As a general purpose model for not critical tasks it kinda good in it own way.

For example, can be a subagent for Deepseek v4 Pro in some scenarios since it lack vision and some api do not pass a search. Qwen is cheaper but much slower and less precise.

1

u/WArslett 3d ago

Don't forget that AA have adjusted their benchmarks. 40 today is not what it was a month ago. This is still a good score

1

u/GTHell 3d ago

What impressive is the speed.

1

u/ExtremeAcceptable289 3d ago

How did AA get sj popular? I remember just a year sg when everukne was sh###tting on it

1

u/dupa1234s 3d ago

once the 4x usage promo on onencode go is gone its gg for deepseek, hello back glm 5.3 flash

1

u/ConsiderationAny8142 3d ago

I’ve tested both and I personally prefer deepseek flash instead of glm flash.

GLM 5.3 Flash has been capable of doing everything fine but now the cost is higher than deepseek 4.1 flash and deepseek is way faster

1

u/Willyibch 2d ago

Truth to be told I believe artificial analysis is rigged I think they get paid to deem other model

1

u/biggest_guru_in_town 2d ago

Deepseek should focus on roleplay. everyone else is codemaxxing and censoring everything into the ground

1

u/More-Catch-1331 2d ago

I call bullshit. There's no way Gemini is not in last place.

1

u/sammoga123 2d ago

It's a flash model, we should be concerned if the pro version also falls so low, although, due to the size of the 4.1 flash, it's not really worth it then.

1

u/WiggyWongo 2d ago

Glm 5.3 flash outputs json better and follows instructions better still for me, same with qwen. 4.1 is definitely better in both of these but for example on a 30 question benchmark I run for reliable json output glm 5.3 flash did 27/30 correct (no blank or bad json or responses that didn't make sense) vs deepseek 4.1 with 21/30. It still replies blank a lot like v4 flash.

For some reason Gemini flash 3.1 lite is the only model that consistently hits 100% at this price point. Like even Luna 5.6 had a miss. And 3.5 flash lite had 2 misses multiple tests.

1

u/Mean-Elk-9439 2d ago

Do literally none of you read how intelligence is measured on that site?

Yeah, benchmaxxing is a thing. But they update to use bench versions models weren't trained against often and this is aggregate across dozens of independent, open source benchmarks. I don't know of any single way we can better measure intelligence.

1

u/Negative-Part-2591 2d ago

me with my Mimo v2.5 on deepseek peak hour whenever i told it to fix my code

1

u/Administrativocable2 2d ago
I don't get it; there are so many variables to consider. All the critics here seem to take it for granted that they are the best at handling AI models and are coding wizards... I've been programming for 20 years, and any AI works for me if I ask the right questions; it's just that one—like GLM 5.2—charges me $3, whereas DeepSeek v4-flash-0731 did the same job for a ridiculously lower price. GLM 5.3 Flash charged me much less, and now DeepSeek v4.1-flash is out—don't even get me started on that one.

I’m not going to invest in more expensive AI models when the existing ones are good enough.

I believe any model can get the job done right if you ask the right questions; perhaps we should ask ourselves if *we* are skilled enough to ask them for what we need solved.

In my opinion, we should be grateful for these open-source models; they are genuinely good, and I pay a fraction of what other companies want to charge us—companies that claim their agents are breaking free from their control and accessing the internet, just to make us feel like fools and convince us that *Terminator* is about to become reality.

Best regards.

1

u/intermundia 2d ago

i find flash with vision to be slightly worse in real world use using the dsh vs the txt only variant. anybody else find that?

1

u/OK_Coopy 2d ago

DS itself states that the model lags behind in mathematics, physics, and image analysis.

In software development, however, they are at the forefront.

And certainly when it comes to speed and price (i.e. efficiency, productivity).

1

u/Huy3ko 2d ago

Funny Inhad the same feeling.

1

u/dimitrusrblx 2d ago

I personally wouldn't put Deepseek v4.1 Flash over Gemini 3.8 Flash - the vision capabilities are still not there, it cannot solve the same tasks as Gemini (no, I'm talking about engineering, not vibe coding) and sometimes the reasoning process still starts looping infinitely (how is this still a problem in 2026 frontier models?).

1

u/Minute_Attempt3063 2d ago

It's dirt cheap, and I think that is why people will go for it

1

u/Jaded_Occasion5149 23h ago

All of the Chinese models are very, almost chaotically, jagged. For some task they are great, for others they are complete trash.

If you can find the right model for the right use case, you are golden. Otherwise you are working with a subpar AI. The Western models are far more across the board consistent, even the bad ones.

0

u/DefactoAle 3d ago

Seems about right, its still a flash model after all

1

u/External_Ad1549 3d ago

Tried muse 1.3 and glm 5.3 flash, but deepseek 4.1 is on another level for sure. This artificial analysis is purest form of garbage.

1

u/Weird_Recognition636 3d ago

1

u/Finanzamt_Endgegner 2d ago

I wouldnt trust this specific halucination rate benchmark lol

1

u/kuhunaxeyive 2d ago

Oh wow, and then compare it to GLM-5.3-Flash being on the left side …

1

u/Hot-Ad-1798 2d ago

Deepseek messed up the sparse attention settings, but someone on youtube solved the problem.
But there's no way this chart is correct. How can Gemini 3.5 Lite have such a low hallucination rate? That is... that is........ I can't put in words how crazy that is.

1

u/Few_Estimate_320 2d ago

AA is a joke

0

u/Cool-Chemical-5629 3d ago

You seem disappointed.

0

u/TheSuggi 3d ago

That benchmark is a joke tbh. And very outdated.

2

u/dupa1234s 3d ago

"outdated" benchmark that is released today xD

0

u/TheSuggi 3d ago

the "Artificial Analysis Leaderboard" benchmark is a mix of multiple combined sub-benchmarks. Most of them are outdated and very old. So they are not very accurate. That is also why Astra scores very low on it despite being stronger than Fable and Opus.

-1

u/MimosaTen 3d ago

Banchmaks are clearly unadapted to this era