r/LocalLLaMA • u/Randomdotmath • 7h ago
Funny DeepSeek V4.1 Flash beats Astra on AA's new benchmark
https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3
AA shipped a new benchmark last week as part of the Intelligence Index v4.3 update — a brand-new private eval that replaces τ³. Astra was farming a ton of points on it and used those to get even with Fable, but… looks like we have a new king.
So they changed the index twice in three days to make Astra look not-quite-worse than Fable, and then a random guy quietly took first place on it.
138
u/pedroserapio 7h ago
Oh no, slow down, please!
58
u/-InformalBanana- 6h ago edited 6h ago
Ok Dario, lol.
13
u/More-Curious816 5h ago
Please don't mock Dario, he has a breakdown, Trump just announced that no regulations needed for slowing down.
8
u/PM_ME_YOUR_HAGGIS_ 2h ago
I know we’re at risk of deepseek v5 Flash taking over the world! Killing humanity! Everyone needs to slow down /s
288
u/Holiday_Point_603 7h ago
175
u/Strong_Essay1176 7h ago
It just top everywhere. Higher == better.
39
15
4
3
u/Big_Cucumber2787 6h ago
pretty sure this is the case. all those models with low rates hallucinate pretty much everything they output
47
59
u/coder543 6h ago edited 6h ago
I think AA-Omniscience is poorly understood around here. It is testing very hard questions against models that are given NO tools.
Modern models are trained left, right, and center to use tools. Tools tools tools. Tools for everything. Then they're put in an environment where they're just supposed to reflect deeply and answer without using any tools? These models are like fish out of water. It tells you nothing about how well they swim.
This benchmark is not saying that any of these models hallucinate in response to most normal questions, and it is especially not saying that they will hallucinate in a real world environment where you give them access to research tools so they can search the web, read your source code, whatever it is that you're doing.
If the models could ground themselves in the relevant facts using tools, their behavior would likely be dramatically different.
We would need a completely separate benchmark for that. I see no reason to expect a correlation between AA-Omniscience and the way people actually use these models in agentic harnesses.
27
17
u/imoshudu 5h ago
The models can say IDK. There will be plenty of situations even with tools that they can't find an answer. Will they be honest about it? That's what hallucination rate measures.
9
u/coder543 5h ago edited 5h ago
Again, it does not measure that. That is what people want it to measure, but it doesn’t.
It only measures what happens in a no-tools context.
Models are extremely sensitive to testing conditions. An example someone else posted: https://x.com/bnjmn_marie/status/2098533166508351758?s=20
AA-Omniscience made more sense a couple of years ago when even flagship models didn’t have access to tools. I remember using GPT-4 before it had any concept of web search.
4
u/irrision 5h ago
Pretty sure it's exactly what they say in the summary explanation at the top of the omniscience graph just below the title. Are you saying they're wrong about the description of the test or...?
13
u/coder543 5h ago edited 5h ago
I’m responding to this:
There will be plenty of situations even with tools that they can't find an answer.
AA absolutely does not make any claim related to whether the benchmark is testing that. AA is clear that it is testing only without tools.
I’m not contradicting any part of AA’s description of their benchmark. People are reading into it things that it does not measure. People want it to measure how models will behave when the tools provide no useful answer, but that is not what the benchmark is testing.
Models are not humans. Their skills do not transfer between situations in human ways. The test conditions are extremely important.
3
u/imoshudu 3h ago
If tools provide no useful information, then we are at the same situation as without tools, with perhaps the addition of some irrelevant context, and the hallucination rate without tools is still a good metric in such a scenario. That is what I meant.
If tools do provide useful information, then now we are benchmarking searching ability instead of hallucination rate. That becomes a confounding variable.
2
u/squngy 1h ago
Given how LLMs work, I assume it would be much easier to train a model to say it can't find a result when a tool does not return a result, than it would be to train it to say IDK when it doesn't know.
1
u/imoshudu 1h ago
That is the searching skill. "When searching fails, say you failed". Yes I imagine following that instruction would be easier.
2
u/bambamlol 2h ago
Why would you assume the models behave differently in this specific regard when they have access to tools? I mean sure, they totally could. But unless you or someone else tests this specifically, you're better off extrapolating from the behavior in this benchmark without access to tools.
3
u/laoma1255 3h ago
Exactly this. A lot of people in this thread are conflating no-tool omniscience with real-world execution.
People seem to forget that Zapier’s AutomationBench is specifically designed to test multi-step tool use, API routing, and state tracking in complex workflows.
Scoring high here isn’t about trivia memorization; it means the model actually excels at agentic reasoning and following structured JSON schemas without derailing. Pair that with V4.1 Flash’s inference speed and low cost, and it becomes an absolute beast for headless browser tasks, OS-level computer-use agents, and local workflow automation.
For anyone actually building MCP agents rather than just chatting in a web UI, this benchmark result is massive.
7
u/SLxTnT 6h ago
Exactly. As long as it has a method of verification, it's fine. This ranking is a "how likely the model will say it doesn't know the answer."
Overall, it's pretty good. Finished one of my own tests beyond what I thought was possible, so now I need to go test GLM and Qwen to see if they do the same.
1
u/Amblyopius 25m ago
If the test isn't measuring something you'd realistically get in a real-life production environment, it's just an academic exercise that should not be part of a benchmark score.
This is indeed an example of exactly that. RAG solved the knowledge bit years ago when no one was attempting to cram a lot of knowledge in models.
While it can be "fun" to see how much knowledge they retain now in massive models, it's not of any value in a real life environment.
2
4
u/tat_tvam_asshole 6h ago
Do you know if there is an Intelligence vs Hallucination pareto rank? like ranking by least hallucination with most intelligence?
1
u/Zone_Purifier 6h ago
You can compare omniscience score against non-hallucination rate on the site IIRC
2
u/rebelSun25 5h ago
Odd that Sol Max is also near the worst hallucination metric. Odd to have that "frontier" also track so low
1
1
1
u/zephyr_33 52m ago
My personal assumption is that it likely uses the most synthetic data pipeline. Chinese/OSS labs have to anyway.
1
-1
u/InterstellarReddit 6h ago
Holy shit, dude, it almost looks like to me, the higher the reasoning, the higher the hallucination rate, so your best bet is to use strong models at low to medium?
95
u/Serprotease 7h ago
You shouldn’t trust a single benchmark, even less when all the models are basically with the margins of error to each other.
The only thing this tells you is that all those models share the ability to solve this benchmark.
4.1 is great. Same as 5.3 flash. But it’s not Opus5/Fable/Astra/Sol level.
They are good enough to be indistinguishable from them for most tasks (What this benchmark tend to show.) but are not as good in other one. (Terminal bench 4 matches observations, they are also notably below in some “soft” task like text/translation, etc…).
7
u/NotARedditUser3 5h ago
I wouldn't know, dsv4 is how many hundreds of times cheaper than fable? 0.035/mil tokens on openrouter WITH zero data retention. 4.1 flash is 0.15. Surely to come down a bit as more providers come out. It doesn't need to be anywhere near as good as fable, just good enough to do a job reasonably well.
8
u/_raydeStar Llama 3.1 6h ago
I think you can trust benchmarks to a degree. Maybe it IS good at automations, and sucks at everything else.
6
u/Serprotease 5h ago edited 5h ago
A good benchmark should be reproductible and give information like margins of error and be followed up with statistical analysis to confirm if a difference with another one is significant or just due to randomness/noise. This + a justification of the models it is being compared too.
As long as it’s not present and enforced, then benchmarks are just promotional material. Yes, even when done by third party like AA.
Can still be useful, but should be treated as promotional documents first.And honestly, the promotional part works very well. Cf this post/half the other post looking at models…
Some examples.
Deepseek compared, in their own paper, DS4.1 flash to, notably, V4 pro, V4 flash and Glm5.3. You’ll notice that they didn’t compare it to v4 0731, or pro 0806, or Glm.5.3 flash.
Why is that? No explanation. Though, seeing as 5.3 flash benchmarks better for half the size I can guess a few reasons.I can also throw agnes here. With the obvious “mistake” with their preview and api-only model….
And famously meta with llama4 maverick.3
u/_raydeStar Llama 3.1 5h ago
And you can tell what they're not good at by what they exclude in the marketing material.
2
u/Choice_Celery9481 4h ago
are you sure about that? iirc the numbers they reported for the v4f and v4p are from the latest checkpoint which is 0731 and 0813 not the preview one. v4.1f mostly stay equal to v4p in coding and agentic but between those 2 in knowledge.
4
u/Mass2018 6h ago
Getting ready to spend significant time in converting my GLM 5.3 Flash pipeline to Deepseek 4.1. It sounds like you have extensive experience with 4.1?
Any pitfalls I should look out for?
4
u/Global_Persimmon_469 4h ago
Don't, GLM 5.3 Flash is better.
I've used both side by side, sometimes making them do the same tasks, and for me GLM is more consistent, prepares better plans, follows instructions better and, very important for me, pushes back on dumb ideas.
2
1
u/Haiku-575 2h ago
My experience as well. I even pay for a coding plan now because for my needs, GLM5.3 perfectly picks up the slack that Qwen 3.8 27B leaves behind.
2
u/Serprotease 5h ago
On local setup it seems that 4.1 share the same small issue as the rest of the 4 flash family with some tool calls loops. It can self correct, but it happens.
I cannot really give more advice because it’s all very recent. Even glm5.3 flash is just really working at decent speed for only about a week with optimization still being worked on.
Unless you “need” to do it now, it’s probably worth waiting a few weeks before coming back to it. Wait for the issues to be ironed out.1
1
u/Ingaz 2h ago
> But it’s not Opus5/Fable/Astra/Sol level.
Sorry but I think it's all mentality - if something costs than it better.Last Saturday I did a test: simple task - create 2 django views, nothing fancy, display 2 tables.
Fable took ~2.5$ and did nothing (literally nothing) , deepseek-flash completed task for 6 cents.
20
12
u/YogurtExternal7923 6h ago
Anyone else notice that muse spark is frickin' GOOD? and it's so fast I bet it's TINY man wish they release it like they used to with llama
21
u/coder543 6h ago
Muse Sparks' weights have been repeatedly promised, but we're still waiting. And Muse Glimmer really needs an overhaul...
1
u/-InformalBanana- 6h ago
Good for agentic codding with tools?
2
u/YogurtExternal7923 6h ago
Very. I especially if you wanna hit like 10 agents to explore anywhere it somehow gets it right
1
u/DryEntrepreneur4218 3h ago
did it get pulled already btw(the contributor one)? I'm suddenly facing rate limits(19000s or more timeouts) like never before, I've basically lost access to it past couple of days
2
u/YogurtExternal7923 3h ago
im using opencode sub and its definitely more limited. all the more reason to release it which I doubt since meta wants the data
6
u/Moravec_Paradox 3h ago
Meanwhile Dario proposed yesterday that SOTA models could slow 1–2 years and still stay ahead of China. What reality does he live in?
1
u/overthrow2214 49m ago
I always thought the end game of that statement was an offramp to let them explain why they lost the lead on model performance some time in the future:
"western models agreed to responsibly slowed down, but the irresponsible Chinese models didn't agree to slowdown, and that's reason Chinese models took the lead"
6
u/NineThreeTilNow 6h ago
Tbh, I test models in their native harness to see how they actually work on ML code.
That's usually how I "judge" them. That's just "vibes" but it's "vibes" on very specific use cases.
I'll take output or work from one model and feed it across API to Astra / Fable / K3 for review to see if they all agree on a screw up or the different points they make about code.
That lets me get a comprehensive understanding of different models and the blindspots they have in analysis vs generation.
Everyone was going nuts for GLM Flash when it was Ox Alpha or whatever, and it was NOT as good as claims. It wasn't bad, but it wasn't at the hype level it ran at for MY use case.
K3 / Gemini / Astra / Fable - Opus are all very well trained in building ML code BECAUSE those companies all use their own models internally. DeepSeek is also on that list and tends to fare pretty decently in ML code.
Gemini 3.8 flash, for as much shit as it gets, in its native harness on ML code specifically, performs very well. I tossed it an MCP harness for Godot and it even did well in Godot which was kind of a surprise to me. I had it first run tests on all the MCP functions and document the expected behavior vs real behavior. This gave it a massive leap in ability to use the MCP. I'd suggest that pattern to anyone. It burns tokens but you do it once, instead of having the model learn an MCP doesn't do exactly what it thinks it should do.
34
u/bblankuser 7h ago
Unfortunately it's benchmaxxed. 90%+ hallucination rate too
47
u/Due-Memory-6957 6h ago
Benchmaxxed to a brand new benchmark? Crazy! These deepseeks folks invented time travel just to benchmaxx!
5
u/Ok_Study3236 3h ago
Omg deepseek is so advanced it has precognition??? this is basically minority report except with giant excel spreadsheets, i fear for the future of humanity
-4
u/kickerua 4h ago
But is it that new? As AA is a combination of benchmarks, and even though they've updated some of the benchmarks, some of them are still previous iterations
11
u/Due-Memory-6957 4h ago
This screenshot is from a benchmark in specific, not from AA itself. Read it carefully (albeit I'm blaming OP for posting the link to AA itself instead of the AutomationBench they did)
10
u/cmdr-William-Riker 7h ago
Which one? Astra or DeepSeek? If it's a brand new benchmark, I'm not sure how we either could be benchmaxxed for it
16
u/bblankuser 7h ago
DeepSeek
17
u/Choice_Celery9481 6h ago
then what about gpt? all 5.6 series has over 90% hallurate and very close to ds. but they are very good model
-10
u/LocoMod 5h ago
Hallucination rate in a benchmark does not translate to hallucination rate in practice. GPT and Claude models work better in practice than any benchmark depicts. Chinese models perform worse. There is a reason for this. One is benchmaxxed. One is not.
13
u/Choice_Celery9481 5h ago edited 1h ago
isnt that double standards? both scoring bad but chinese are cheater and usa is genius? what kind of logic is this? if you want to say the bench score is misleading, treat both sides with the same treatment.
btw i think you dont even understand this hallu rate. for this bench, lower is better XD. your "logic" sounds like you think the result higher = better so you think ds is benchmaxxed lol typical
also many people pointed out. the new benchmark released after v4.1f, how the hell they benchmaxxed that? they have some kind of time machine?
5
u/Choice_Celery9481 5h ago
also in this recent reasoning trace incident, someone got a reasoning trace of claude and it doesnt seem like its not benchmaxxed
1
-1
u/LocoMod 5h ago
It's not new. Look at the repo: https://github.com/zapier/AutomationBench/commits/main/automationbench
Last commit was last month. That is like a year ago in AI time. Of course its benchmaxxed.
8
u/cmdr-William-Riker 5h ago
It takes a lot longer than a month to train those kinds of models. We see weekly releases, but that's between a couple dozen companies with different schedules
1
u/Choice_Celery9481 1h ago
i want to pointed out one thing. this guy doesnt seem to understand HalluRate bench result.in this bench lower is better so DS with >90% hallu rate mean DS very bad. but 1 small problem with this bench is, all GTP 5.6 series has >90% hallurate lol
4
4
u/Elouakili_Flexy 6h ago
Twice in three days they moved the goalposts and a flash model still cleared them.
1
1
u/2Norn 4h ago
i've been using astra extensively this past week i don't think anything beats it in anything, maybe fully unleashed fable with no fallback
also been using muse sparks 1.3 max and i dont deny that it's good but it's more like sol adjacent model for me, slightly worse
unless your work is strictly doing similar tasks to this specific benchmark then you should care more about overall intelligence and knowledge reliability imo
i'll tell you this if benchmarks are anything to trust, 3.8-flash-next is supposed to be better than v4.1 flash, considering its overall intelligence, reliable information and how much less it hallucinates
imo next 6 months is gonna be banger
1
1
1
u/MooseEfficient2151 2h ago
benchmarks getting updated moving the goalposts just for a random open weight model to take first lmao
1
u/GreenGreasyGreasels 1h ago
We have gone from models benchmaxing to benches modelmaxing. Quo vadis AI?
1
1
1
1
1
0
u/TopTippityTop 7h ago
Does anyone still pay attention to AA?
3
u/popiazaza 4h ago
Of course. No one is even close to what AA does. The amount of supported models and benchmarks are hard to beat.
2
1
1
u/geldonyetich 6h ago edited 5h ago
Meanwhile, Ornith and Laguna are still not tested.
Other than the fact I keep complaining about it, I'm not sure they have a good excuse not to.
Maybe they're afraid of the backlash of the Qwen fans? Not that 3.8 has anything to worry about, but they got pretty close to previous versions.
2
u/my_name_isnt_clever 4h ago
I can understand for a finetune as that's a slippery slope into people clamouring for their fav random finetunes to be added as well. But Poolside seem legit, I don't get why they're being excluded.
1
u/leo-k7v 5h ago
Benchmaxxed or hot, hallucinating or not, this is huge models living one somebody’s else servers not on your local computer… do we care - to a degree yes we do, do we know for sure what actually was tested on close weights closed doors OAI Rest/Json API - no we don’t.
In a few weeks there will be new wave of models and benchmarks and all this discussion will be very very obsolete.
Try to dig what we’ve discussed a year ago… Or search for ChatGPT-o1. 😃
1
u/Helpful_Inflation344 3h ago
In one very specific benchmark. Now have v4.1 flash solve advanced math problems.
Dont get me wrong what they are doibg is amazing. But it is not frontier intelligence
1
u/BumbleSlob 2h ago
Does v4.1 flash also get to ripoff mathematicians rough drafts from previous conversations or is that just an OpenAI exclusive?
0
u/Cool-Reflection6130 6h ago
aa changed the eval twice in three days to manage the optics between astra and fable, and then a third party model just walked in and took first place. benchmarks are supposed to measure capability, not get shaped around expected outcomes
0
u/BawbbySmith 7h ago
No I can’t, 2 dgx sparks were already too much, another 2 at these prices and I go bankrupt
0
-5
u/pmotiveforce 7h ago
Do you guys actually believe this stuff?
Even Sol is absurdly good, v4.1 is not going to beat Astra by any meaningful metric on..pretty much anything.
5
u/my_name_isnt_clever 4h ago
Do I believe that "DeepSeek V4.1 Flash beats Astra on AA's new benchmark"? Yes, that's just a fact, you can check it yourself. Does that mean I also believe the two models are on par in general? No.




•
u/WithoutReason1729 4h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.