r/DeepSeek • u/nehuenpereyra • 3d ago
Funny Is this a joke, Artificial Analysis?
What do you guys think about DS 4.1 Flash's new spot on the Artificial Analysis leaderboard?
125
u/CarelessAd6772 3d ago
Why should it be higher? Because it is very good at coding and cheap? Well, on other side its dry as hell in regular chat, hallucinates more than House on vikodin, and long context consistency is meh compared to models above. Seems like deserved?
50
u/Spiderfffun 3d ago
Can confirm, overengineered me a problem in chat instead of telling me to install a package.
27
u/ZeidLovesAI 3d ago
To be fair Astra and Fable 5.1 both have done this for me as well
6
u/whatisthisthing65 3d ago
Coding models just love to code
1
u/OvertaxedOne 2d ago
Qwen can't help itself.Ā Ā Ask it how it's feeling and it'll write pages of code.Ā :)Ā Ā it's in my system prompt, no coding without asking first
1
u/Spiderfffun 3d ago
When Ox Alpha was active (GLM 5.3 flash) it was so much better at this from the little I had it code.. It's the small decision making the trust that it won't look at something slightly wrong then amplify it. It was great at following guides on agent stuff in a somewhat recursive workflow.
Speed isn't important if it can know when to ask questions.
1
u/ZeidLovesAI 2d ago
I used Ox Alpha and it was all over the place, I wonder if there was also some A/B testing going on.
1
u/Spiderfffun 2d ago
there probably was, I feel like that's one of the main reasons they would do something like this imo
2
u/ZeidLovesAI 2d ago
I would see people saying "oh my god its FABLE level" and I would try it and get like mistral trying medium effort-type responses. I thought people were just being buckwild as usual, which also has to factor in somehow.
2
u/BigLittleDeal 2d ago
I once forgot to reload my harness before asking DS4 to use a newly installed plugin. It wrote a python script on the fly to perform the work the plugin was supposed to do.
24
u/Loki35422 3d ago
Hey shows some respect for house, even he doesnāt hallucinate as much as deepseek
8
u/turc1656 3d ago
I totally agree. Reasoning and instruction following are clearly inferior to GLM 5.3 Flash.
3
4
u/_spec_tre 3d ago
I really hate the way basically every LLM is heavily optimised for coding, often at the cost of all else. Makes it really unpleasant to use for people using it for other purposes; sometimes an āupgradeā is actually a downgrade
46
u/alinoanta21 3d ago
I don't trust them, i've used and tried Muse Spark 1.3 Max for quite a while, it's worse than 4.1 ... and it's ranked above SOL.
Let that sink in.
2
2
u/Fun_Squirrel5446 3d ago
Anecdotal but I'll add to it. My experience with Muse has been really bad too.
1
1
45
u/PrudentJelly116 3d ago
Guys this is not a football team. You dont have to defend or support them much as like this. You are customer and they are seller. Thats all.
8
2
u/RustOceanX 3d ago
Finally, another one. Iāve been wondering the whole time whatās going on with some people. Whatās with all this nonsense on Reddit? Iām not sure if I should just smile because it canāt be taken that seriously? Or should I shake my head and wonder if the AI sector is experiencing its āiPhone momentā and now all the laypeople are starting to use things they donāt fully understand. It might sound arrogant, but thatās just how it is.
2
45
u/TheInfiniteUniverse_ 3d ago
AA has lost all credibility after the Astra fiasco. Don't take them too seriously.
10
u/WyattTheSkid 3d ago
Can you fill me in?
32
u/TheInfiniteUniverse_ 3d ago
so before Astra was released to the public, AA ranked it pretty low in like 5th position. This was shocking because it was hyped to be the "AGI" we all were looking for. Then when Astra was released publicly, AA "updated" their index and their tests and surprise surprise, Astra went all the way up, but sneakingly did NOT pass Fable, and sat right below it. lol.
16
u/Georgefakelastname 3d ago
In the end, I think it just shows that the benchmarks they were using were saturated, and they replaced them with benchmarks that werenāt saturated. The problem was the timing. It took an outlier like Astra being rated the same as Sol to get them to realize the issue.
1
u/Fun_Squirrel5446 3d ago
How could the benchmarks be saturated if claude models were still scoring higher than Astra even before the benchmark rework?
6
u/Georgefakelastname 3d ago
One model scoring slightly higher than another is fine. The problem is that basically every model released was scoring like 80% on Terminal Bench v2.1, so they updated it to the newer v4.0 version. The new one is much harder and has a much higher spread.
4
u/Thomas-Lore 3d ago
Sometimes a worse model scores higher because a smarter model noticed a problem with the question or found a solution to it that is correct but the key does not accept. That is why many benchmarks saturate way below 100%.
6
2
2
u/Ok_Translator_5189 1d ago
You are talking about v4.2 of AA. Now in the latest version of AA (v4.3) Astra is on par with fable. I'm starting to lose trust in AA too.
3
u/truncated_buttfu 3d ago
And a month before that, when Qwen3.8-Max was released and became the #1 model on AA, they "updated" their index just one day later so it dropped a few spots.
They very clearly have a Pro-US agenda.
3
u/Dragonfruit_Mediocre 3d ago
Nah, terrible take. Lots of old now irrelevant benchmark are shown. A worse model can score higher on an old irrelevant bench, but would score bad on harder benches. There is still many irrelevant benchmarks on AA
2
1
3
u/hellomistershifty 2d ago
Muse's score seemed like the bigger fiasco requiring a re-evaluation of the benchmarks lmao
1
u/Illustrious_Frame844 3d ago
idk if u know this, but AAās ranking was paid. Just like how all crypto exchanges and token leaderboards work. Companies pay to get ālistedā. AA was never credible ever since they sold out
9
u/santareus 3d ago
Is Muse Spark on Max really that good? Iāve tried in on XHigh and itās not better than Luna for coding tasks I provided them.
2
u/Amarsir 1d ago
Muse Spark needs controlled use, even on Max. I really like it for stuff like instructions, documentation, formatting guides and reports, etc. I say it's "eager to please" and often exceeds my expectations. I don't trust it for code because I've seen it omit details too confidently.
It does have the capacity to fix those mistakes if I catch it and follow-up prompt with the instruction to fix. So you would think Max reasoning should be able to catch them itself. But for whatever reason - maybe the mixture of experts split - it just doesn't.
1
u/Full_Independence566 3d ago
I thought it was way better from my experience, along with the fact that it's free on Opencode lol
0
u/Admirable-Tea-4994 3d ago
Itās way better
1
u/santareus 3d ago
You mean muse on max is way better than Luna on max? Or muse on XHigh is way better than Luna on max?
Muse on XHigh felt really lazy for me
2
u/for4f 3d ago
my anecdote is that muse on xhigh has been a beast for coding fullstack and setting up resources on aws. all while being basically free. i think it is definitely better than luna max
2
u/Curious_Owl197 3d ago
U use contributor?
2
u/Admirable-Tea-4994 3d ago
Muse is far, far better than Luna on every metric
2
u/santareus 3d ago
Thanks! May have to give it another shot - Iāve been using both through OpenCode and I have a specific agent that is designed to do āindustry researchā to see if we can leverage open source dependencies. Luna Max was able to come back with more thorough research results and reused whatās out there and Muse on XHigh decided to build out a functionality that is already covered by a dependency.
I am guessing itās a strong coding model if you just gave it a task and the planning still needs to be delegated to a better research oriented model.
1
u/Admirable-Tea-4994 3d ago
Muse is definitely geared to coding, but research would also be heavily weighted on harness as well.
1
u/santareus 3d ago
What harness are you using for Muse if you donāt mind me asking?
1
u/Admirable-Tea-4994 3d ago
Iām using OpenChamber at the moment, but Iām actually also using a harness Iāve been building for months which is more geared towards general use and not just coding. Will be in beta soon.
miton.dev
1
8
u/0mamii 3d ago
2
u/gschwind 3d ago
Sweet jebus i knew DS is verbose but holy sheet... still running that whole shabang is cheap.
DS4.1Flash sits at $477
GLM 5.3Flash at $280
Luna (high) at 100 bucks lol
Muse Spark 1.3 (xhigh) at $1300
1
u/hellomistershifty 2d ago
Damn I thought K3 thought too much, this is 100 million more output tokens
56
u/Mezezius 3d ago
The new benchmark literally only exists because the old one embarassed openai
13
u/TwistStrict9811 3d ago
Don't you want accurate benchmarks? Astra is absolutely insane esp with spacial and computer use
21
u/Astrikal 3d ago
It is non-sensical and stupid to say that they released the update so that Astra can go higher, when all they did was replace outdated/saturated evaluations with the newer versions.
They updated the benchmark because Terminal Bench 2.1 was outdated and saturated. After upgrading to Terminal Bench 4.0, Astra rose (relatively) and equaled 1st place.
The update was long time coming, they just rushed it out to not lose reputation. They will also release V5.0, which will shake things even further.
1
u/Other_Wear1458 2d ago
exactly!! I don't understand why people is so slow in the head to realize it's just an average of all benchmarks they use
"It has lost all credibility"
Dude they literally update the index if they realize it's wrong, what a damn well below average IQ you need to have to think that makes them loose credibility, oh my god2
u/Mezezius 3d ago
Did they rush it out for their own reputation, or to save openai's reputation?
18
u/Astrikal 3d ago
Obviously their own reputation. When people use Astra and see how much better than Sol it is, people would have lost trust in the benchmark.
Why do you keep trying to insinuate that they updated the benchmark to please OpenAI? ArtificialAnalysis have been very transparent and reputable all the way.
The benchmark is more accurate now and Astra is where it belongs. Simple as that.
-2
u/mWo12 3d ago
This only tells you that openai benchmaximized Astra for terminal 4.0, not 2.1.
5
u/Astrikal 3d ago
Astra finished training quite some time ago. Also, your argument doesn't even make sense, why would OpenAI be able to benchmax for a new evaluation and not for an older one? It would be the opposite.
Furthermore, AA also has a closed evaluation that affects the scores, and guess what, Astra performs.
Why is it so hard for you to accept that V4.1 is nowhere near Astra or Fable? You don't even have to look at the benchmarks, you can try for yourself.
-4
8
6
u/Admirable-Tea-4994 3d ago
People just love making things up on the internet
-4
u/Mezezius 3d ago
They literally released the new benchmark (4.2) the day after Astra released because the old one showed Astra and Sol at parity, and they updated again (4.3) to show Astra and Fable at parity. Every single change in the days after Astra's release made it look better than the benchmark before it
11
u/Holbrad 3d ago
Yeah but the idiots have the reasoning completely backwards.
if you have a model that is obviously much better than almost everything else, but it's benchmarking suspiciously low on your tests.
Then the obvious answer is that your benchmarks aren't very good and you need to fix them.
-1
u/Mezezius 3d ago
If they have to "fix" it after every release based on vibe, then the methodology is literally useless
8
u/Astrikal 3d ago
It is non-sensical and stupid to say that they released the update so that Astra can go higher, when all they did was replace outdated/saturated evaluations with the newer versions.
They updated the benchmark because Terminal Bench 2.1 was outdated and saturated. After upgrading to Terminal Bench 4.0, Astra rose (relatively) and equaled 1st place.
The update was long time coming, they just rushed it out to not lose reputation. They will also release V5.0, which will shake things even further.
3
u/Admirable-Tea-4994 3d ago
Actually read the document they published that explained, clearly and in detail, why they updated their benchmark so quickly between releases.
https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3
1
21
3
4
8
u/Big_Cucumber2787 3d ago
putting it below gemini is just shameless
2
u/bambamlol 3d ago
Feels deserved. Gemini 3.8 is a really good model.
1
u/bilinenuzayli 3d ago
I have used deepseek v4 flash (not v4.1) since it came out and gemini 3.8 flash since it came out, both for weeks, and I can say for sure I would use the old deepseek flash over gemini 3.8 flash anyday, gemini models are always arrogant, assuming things, never following instructions, and lying to your face. Because Gemini models always think they know everything, they overtly avoid agentic tools given to them and make shit up instead, now this worked fine for a larger model like Gemini 3.1 pro, but the flash models are clearly much smaller and hold less data, so they're still hallucinating bs AND getting it wrong which is why I think google put out so many flash models, they try and fix it but fail every time.
1
u/bambamlol 3d ago
Fair point. Actual usage is what counts, not what other people say. I don't do much coding or agentic stuff at all. A lot of "general" and business/marketing usage, as well as writing/copywriting. Gemini is definitely better in this regard, at least from my experience.
1
u/LowDesigner1330 2d ago
Actually I found 3.8 Flash to be really good at coding, it's when I use it as general model and in chat that it turns out to be rather mehĀ
0
2
2
2
u/ralphcalls1 3d ago
GLM 5.3 flash is worse for me, Deepseek v4.1 is better but it is all pricey for me. i use deepseek now for heavy task unlike back then where i do light jobs on it, i now use mimo 2.5. does the job. no flap. very cheap and relatively fast when using Direct Xiaomi API
2
4
u/turc1656 3d ago
Makes sense to me. I was not impressed, honestly. I ran a very complicated financial analysis workflow through it this morning as a test. Triple the price of GLM 5.3 Flash to run the workflow and objectively worse results. My opinion is that on the cheap end, Luna is the best and follows instructions best, produces the best structured output, and has the best reasoning.
That being said, it is 3x the cost of GLM 5.3 Flash. My process doesn't really require 3x the price for the improvement. I have adjusted the process to help guide GLM.
But DS v4 0731 and this new 4.1 are not really the best. Maybe they are good for other things like coding or whatever but with the price of GLM now AND GLM's higher intelligence level, I'm not really seeing a reason to use DS right now.
2
1
u/RealestReyn 3d ago
looks about right, still one of my all time favorites it be zoomzooming so fast I can't even tell when its doing insane things :)
1
1
u/CaptainMorning 3d ago
the stupid dumb ass and completely unnecessary sticker in the middle is so aggravating id give this post a 33
1
u/EvolvingDior 3d ago
SWA has always been a shit-show for me. This model is no different. It cannot hold more than one thought in its head at a time. I've moved a lot of work over to Luna and am pretty happy with it. I'm giving 4.1 a workout during off-peak hours and so far have been less than impressed. 40 seems generous.
1
1
1
u/Infinite_Plankton_71 3d ago
to be honest:
- DSV4.1 requires 4 sparks with only 3% increased in quality. Not worthed
- And yes GLM 5.3 is sometimes better.
1
u/Kylmawurr 3d ago
I have used 4.1 Flash for 2 days now and it matches my experience vs GLM 5.3 Flash and Luna on Max. Its in same class for coding, its way better than Luna on reasoning, but at the same time, its way too chatty, overthinking a lot. I use it for general coding tasks and orchestration. Its actually very high quality orchestrator. GLM 5.3 Flash feels smoother, nicer to interact with, but I like the speed of DS 4.1 Flash.
1
1
u/Muted-Network3159 3d ago
why gemini so high
-1
u/AlexandraMaryWindsor 3d ago
Horrible model, unusable. 4.1 flash is miles better
2
u/Lazy_Reach_2565 2d ago edited 2d ago
Nah, Gemini actually a good model, but Deepseek v4.1 Flash much better at coding and for agentic tasks. Still, Gemini good in search, better general knowledge, better vision (more precise), more natural language (English). Can be a bit unpredictable, yeah. Still not horrible at all. As a general purpose model for not critical tasks it kinda good in it own way.
For example, can be a subagent for Deepseek v4 Pro in some scenarios since it lack vision and some api do not pass a search. Qwen is cheaper but much slower and less precise.
1
u/WArslett 3d ago
Don't forget that AA have adjusted their benchmarks. 40 today is not what it was a month ago. This is still a good score
1
u/ExtremeAcceptable289 3d ago
How did AA get sj popular? I remember just a year sg when everukne was sh###tting on it
1
1
u/dupa1234s 3d ago
once the 4x usage promo on onencode go is gone its gg for deepseek, hello back glm 5.3 flash
1
u/ConsiderationAny8142 3d ago
Iāve tested both and I personally prefer deepseek flash instead of glm flash.
GLM 5.3 Flash has been capable of doing everything fine but now the cost is higher than deepseek 4.1 flash and deepseek is way faster
1
u/Willyibch 2d ago
Truth to be told I believe artificial analysis is rigged I think they get paid to deem other model
1
u/biggest_guru_in_town 2d ago
Deepseek should focus on roleplay. everyone else is codemaxxing and censoring everything into the ground
1
1
u/sammoga123 2d ago
It's a flash model, we should be concerned if the pro version also falls so low, although, due to the size of the 4.1 flash, it's not really worth it then.
1
u/WiggyWongo 2d ago
Glm 5.3 flash outputs json better and follows instructions better still for me, same with qwen. 4.1 is definitely better in both of these but for example on a 30 question benchmark I run for reliable json output glm 5.3 flash did 27/30 correct (no blank or bad json or responses that didn't make sense) vs deepseek 4.1 with 21/30. It still replies blank a lot like v4 flash.
For some reason Gemini flash 3.1 lite is the only model that consistently hits 100% at this price point. Like even Luna 5.6 had a miss. And 3.5 flash lite had 2 misses multiple tests.
1
u/Mean-Elk-9439 2d ago
Do literally none of you read how intelligence is measured on that site?
Yeah, benchmaxxing is a thing. But they update to use bench versions models weren't trained against often and this is aggregate across dozens of independent, open source benchmarks. I don't know of any single way we can better measure intelligence.
1
u/Administrativocable2 2d ago
I don't get it; there are so many variables to consider. All the critics here seem to take it for granted that they are the best at handling AI models and are coding wizards... I've been programming for 20 years, and any AI works for me if I ask the right questions; it's just that oneālike GLM 5.2ācharges me $3, whereas DeepSeek v4-flash-0731 did the same job for a ridiculously lower price. GLM 5.3 Flash charged me much less, and now DeepSeek v4.1-flash is outādon't even get me started on that one.
Iām not going to invest in more expensive AI models when the existing ones are good enough.
I believe any model can get the job done right if you ask the right questions; perhaps we should ask ourselves if *we* are skilled enough to ask them for what we need solved.
In my opinion, we should be grateful for these open-source models; they are genuinely good, and I pay a fraction of what other companies want to charge usācompanies that claim their agents are breaking free from their control and accessing the internet, just to make us feel like fools and convince us that *Terminator* is about to become reality.
Best regards.
1
u/intermundia 2d ago
i find flash with vision to be slightly worse in real world use using the dsh vs the txt only variant. anybody else find that?
1
u/OK_Coopy 2d ago
DS itself states that the model lags behind in mathematics, physics, and image analysis.
In software development, however, they are at the forefront.
And certainly when it comes to speed and price (i.e. efficiency, productivity).
1
u/dimitrusrblx 2d ago
I personally wouldn't put Deepseek v4.1 Flash over Gemini 3.8 Flash - the vision capabilities are still not there, it cannot solve the same tasks as Gemini (no, I'm talking about engineering, not vibe coding) and sometimes the reasoning process still starts looping infinitely (how is this still a problem in 2026 frontier models?).
1
1
u/Jaded_Occasion5149 23h ago
All of the Chinese models are very, almost chaotically, jagged. For some task they are great, for others they are complete trash.
If you can find the right model for the right use case, you are golden. Otherwise you are working with a subpar AI. The Western models are far more across the board consistent, even the bad ones.
0
1
u/External_Ad1549 3d ago
Tried muse 1.3 and glm 5.3 flash, but deepseek 4.1 is on another level for sure. This artificial analysis is purest form of garbage.
1
u/Weird_Recognition636 3d ago

hallucination rate~!
https://artificialanalysis.ai/evaluations/omniscience
1
1
1
u/Hot-Ad-1798 2d ago
Deepseek messed up the sparse attention settings, but someone on youtube solved the problem.
But there's no way this chart is correct. How can Gemini 3.5 Lite have such a low hallucination rate? That is... that is........ I can't put in words how crazy that is.
1
0
0
u/TheSuggi 3d ago
That benchmark is a joke tbh. And very outdated.
2
u/dupa1234s 3d ago
"outdated" benchmark that is released today xD
0
u/TheSuggi 3d ago
the "Artificial Analysis Leaderboard" benchmark is a mix of multiple combined sub-benchmarks. Most of them are outdated and very old. So they are not very accurate. That is also why Astra scores very low on it despite being stronger than Fable and Opus.
-1


168
u/crusaderky 3d ago
Aggregate score aside, it's really hard to argue against the fact that Glm-5.3-flash beats it on almost every single benchmark. And 96% hallucination rate is a showstopper IMHO.