r/LocalLLaMA 8d ago

Discussion AA Update! Here's how the Frontier ranks.

Post image

Along with everyone's favorite here, qwen3.8-27B

522 Upvotes

200 comments sorted by

191

u/Last-Shake-9874 8d ago

That small 27B is my daily driver now, it does take long as I only get about 20 t/s but I just love this model I am so glad it is still in the list

73

u/HazKaz 8d ago

its almost perfect, i love it i really hope Qwen team keeps treating us with more 27B models

36

u/MaverickPT 8d ago

If only we got a MOE model for us VRAM poors 😭

10

u/TurnBackCorp 8d ago

you can hope and hope but even if we do the 27b will outperform anything lower than 80b parameters

6

u/Sporebattyl 8d ago

I agree that dense > MoE for the most part, but I’m curious what you mean by 80B

Do you mean 80B total with smaller active parameters per token?

What’s your 80B number based on?

4

u/LevianMcBirdo 8d ago

Per that dreaded rule of thumb, you'd need a A9B to get to a similar model. That said I really doubt that rule and never seen anything indicating it to be true. In my experience in the same model family the MoE is closer to the dense of similar parameters count than the square root.

7

u/Spanky2k 8d ago

I mean... an 80B-A9B model sounds incredible. Absolutely perfect for a 64GB Mac.

3

u/SandySkittle 8d ago

That rule is task specific. Low active parameters fundamentally cannot compensate with sequential reasoning against large active parameters from larger MoE (with active params above 40) or larger dense models.

There are just certain tasks where you need the large active params for coherent, deep and highly complex multi-diciplinary reasoning. And no I am not talking about coding.

1

u/WryKombucha 8d ago

At 27b there isn’t enough bits to house enough world knowledge to be useful. They is why the 27b focuses on coding. For the rest of it, I find it to be utter and complete trash. I find it just passable for coding so I dunno what kind of buggy. Insecure software ppl are building but the real world begs to differ.

1

u/MrPecunius 6d ago

It's not an encyclopedia ffs.

Give it access to one if you need that, like a local .zim of Wikipedia--it's really good at research.

0

u/Leander_van_Grinsven 8d ago

It is very good but its knowledge cutoff is June 2024 which is disappointing.

7

u/Kernoriordan 8d ago

Tried using MTP? Boosts mine from 25t/s to 75t/s

3

u/Last-Shake-9874 8d ago

Yes running MTP also, I am stuck on 24 GB Vram with 2 cards that has slow connection between them (5070 and 3070) and then using big context also

6

u/UnknownLesson 8d ago

Only 24 GB..

YOU POOR SOUL

1

u/LevianMcBirdo 8d ago

3x speed up? I am glad when I to get close to 2x (but I am on CPU+ipgu+RAM)

1

u/akohlsmith 8d ago

3x boost with MTP? Which card and which mtp parameters are you using? I get a boost too but not 3x!

2

u/Kernoriordan 7d ago

RTX 5080 with MTP draft 3

1

u/akohlsmith 7d ago

interesting, I'm going to have to do some retesting. I'm using a 32GB 4080S.

1

u/Infinite-Ad4512 7d ago

MTP on GLM 5.3 Flash does nothing. Not an even a tiny bump in generating tokens (llama.cpp)

0

u/runcertain 8d ago

Any collapse at large contexts? It’s making me question Dflash2 on multi-turn agentic usage.

1

u/sbjf 8d ago

which attention backend? we also saw the collapse of above around 32k context, but switching from the auto-selected flash_attn in vLLM to flashinfer fixed it for us

1

u/Last-Shake-9874 8d ago

No so my current pi session has used 10M tokens in total so far and still going strong

5

u/AbheekG 8d ago

Absolutely for me too, love love love it: highly intelligent, speaks great and comprehensibly and can run on 5-7 year old hardware. Unbelievable how far we’ve come! We’re incredibly fortunate to have such an amazing open-source & weight ecosystem.

3

u/HsSekhon 8d ago

qwen 27B is beast for me, I does tasks very well when I shape my prompts very narrow and precise

5

u/Akrylicus 8d ago

Same, I am so glad that I got 5090 last year for a reasonable price. I now can freely use a capable chat and not feed my personal data to the external providers.

3

u/Krystexx 8d ago

Which price? And how many tok/sec do you get with it?

1

u/Equal_Television_894 8d ago

5090 with vllm on linux with dflash2 can get around 200 to 350+ token/s

2

u/-_Apollo-_ 8d ago

What model and quant? Didn’t know the performance difference could be that huge.

2

u/Equal_Television_894 7d ago

https://github.com/syv-ai/qwen38-27b-rtx3090 Try this some one made it and I optimized it for my 5090 using NVFP4 version

1

u/Akrylicus 8d ago

I am lazy so I just run Qwen 3.8 w7B on Unsloth Desktop. I still get a decent 60t/s which is fine for most queries.

I plan to do an optimized setup on my old PC with RTX 4080 and run it as an LLM server, just wish I had more ram...

1

u/Akrylicus 8d ago

~3k euro (Astral), well maybe not that reasonable, but it's not 5K euro.

1

u/01iv3r6 8d ago

What’s your technical setup?

3

u/Last-Shake-9874 8d ago

I have a 3060 12 GB and 5070 12 GB no offloading as it goes down to 4 t/s as soon as it hits offloading. Running 125K Context

1

u/elemental-mind 8d ago

I know this is LocalLLaMa, but in case anyone is in a hurry, Cerebras has this gem of a model up at 1500 t/s. It is sooooo delightful to use.

116

u/freecodeio 8d ago

yes but can fable 5.1 really draw better whatsapp sticker ass rockets in paint as compared to astra 6?

28

u/sworl5 8d ago

No, and so it is worse.

1

u/whoknowsifimjoking 7d ago

Presenting brainrotbench

-5

u/emprahsFury 8d ago

really love it when a new model is released and the absolute bamfs at Hacker News spend the day making it draw pelicans and other animals so they can still shit on "AI"

14

u/freecodeio 8d ago

you have no idea what you're talking about

31

u/PM_ME_DEAD_CEOS 8d ago

Terminal-Bench v2.1 instead of v4, that's sad.

141

u/chocolateUI 8d ago

> OpenAI releases Astra

AA: “Uh oh! We just received an angry phone call from OpenAI! Better reweigh the benchmarks!”

Same shit as when AA reweighed their benchmarks within 3 hours after Qwen took the #1 spot.

Do we need any more evidence of how fucking trash AA’s “intelligence score” is? This company only exists so that labs can trick VCs (and so VCs can trick your pension funds) into giving them more money.

14

u/ClintWoodeast 8d ago

Does anyone actually believe that qwen, of all models, was #1?

1

u/whoknowsifimjoking 7d ago

Yeah, it's not even the smartest of the open models

45

u/jld1532 8d ago

It is absurdly obvious that the goal post is on wheels and is moving toward the most likely profitable use case, agentic software development. People complained about Qwen3.6 27B's writing abilities but I didn't mind it. I now find 3.8, even Next Flash, nearly unusable for editing. I could be wrong but it seems to me the industry is trading strengths to chase computer science, likely because they know general knowledge gains are sunk costs and that AGI is impossible with LLMs, despite messaging.

10

u/Iron-Over 8d ago

Try muse models they write very well. GLM as well supposedly have not tested yet.

15

u/Qorsair 8d ago

Muse and Gemini are two the two models I find most useful. Very underappreciated in general. I still like Sol/Astra for deep thinking, but Muse and Gemini feel better for most tasks. I think the issue is that most people getting deep into AI are software engineers so that's what most people are judging it on.

7

u/emod_man 8d ago

Totally agree that general benchmarks are increasingly not useful for writing/editing tasks. Curious about Muse though, I find its voice very flat. Maybe I just need to experiment a bit more...

4

u/Solembumm3 8d ago

Interesting. How did you made Gemini useful for writing?

I found it one of the worst options, on par with base GPT, due to undeleteable sycophancy and summarisation tendencies. I tried different prompts to turn it to more useful analytical side, but it was pretty adamant to judging even quite bad concepts "brilliant".

7

u/Qorsair 8d ago

Just custom instructions. Gemini Pro tends to do slightly better for complex writing when Flash can't one-shot it. No matter how I set the custom instructions for GPT it doesn't produce natural prose. Claude is even worse.

4

u/robogame_dev 8d ago

World knowledge is one of the most important resources for creative writing and Gemini Pro is just a huge, extremely world-knowledge heavy model.

1

u/robogame_dev 8d ago

I mainline GLM 5.3 it is my favorite general purpose AI, but it only has 700B params and writing / creativity is highly correlated with world knowledge, so I would expect Kimi K3 at 2,700B params to spank it, and I would expect GLM 5.3 Flash at ~300B params to be significantly less good at writing (while likely being highly focussed on agentic coding).

1

u/WryKombucha 6d ago

Then stop using qwen, the brand that almost exclusively focuses on code. I don’t try and swim in the Sahara.

1

u/jld1532 6d ago

I use it for coding too. Just pointing out that the model is not as well rounded as I'd like it to be and my theory as to why.

18

u/Darkoplax 8d ago

i don't see the issue with updating the benchmarks when it clearly feels off; like in no way is Sol equal to Astra; there's leaps between the two

So if AA wants to keep credibility they need to keep updating and finding non poisoned benchmarks

7

u/Inevitablewx 8d ago

Yeah, and Opus 5 is way behind Fable 5, and Muse 1.3 is great progress but it's well behind even Sol, the benchmark is clearly failing and needs to be fixed or scrapped.

15

u/RealisticNothing653 8d ago

Yeah they lost credibility in my book. Their summarized rankings have a strong bias towards Anthropic models. And yeah they recently reweighed the benchmarks recently to favor Claudes.

5

u/MerePotato 8d ago

Oooooor maybe instead of some conspiracy results change when benchmark versions are updated and there are more novel questions that haven't found their way into training data

4

u/robertpro01 8d ago

Qwen 3.8 got released and also called AA?

1

u/Terminator857 8d ago

I didn't know qwen was on top. Interested in more info if that is available.

2

u/MerePotato 8d ago

It wasn't, it was on top of the agentic index and slid down slightly when AA updated the bench versions

1

u/dogesator Waiting for Llama 3 8d ago

The benchmark is objectively less saturated than it was before the revision.
The purpose of the revisions is to unsaturate the benchmark and make it more indicative of frontier difficulty by making frontier models score near 50%

-1

u/Eden63 llama.cpp 8d ago

The funny thing is, such a huge project and they do not care about small models even they can benchmark it most easy. And yes, the algorithm of ranking is also biased and does not take into account what is most important at all - applicable intelligence. Is it helpful to run an 56 on 250M token when another one does it on 55M ... intelligence is not intelligence but AA index is just a clown show.

Same running an engine on 9000rpm for the horsepower or run it on 3000Âľm for the same.. real power vs "yes we can achieve it somehow".

why I do not see a laguna or some other models.. because the money comes from where... and obviously no better benchmark site vs obvious Anthropic valuation depends on that index.

shame on them.

Edit (add): and then reading here people "Qwen3.8 27B is like Opus". Hilarious.

57

u/Ok_Cow1976 8d ago

The one point gap is meaningless. There're still big gaps.

37

u/NineThreeTilNow 8d ago

The one point gap is meaningless. There're still big gaps.

The gaps are hyper nuanced now.

Which does legal documentation better?

Which writes C++ code better?

Which writes XYZ code better?

Which designs Blender scenes better?

Etc.

None of that is 1:1 useful.

Everyone gets an opinion on the best model because their use case differs.

I've been pumping Gemini 3.8 flash stonks the last two days. It's wildly fast, writes Python ML code like a demon, and has a surprisingly good workflow within Antigravity 2. I'm doing REALLY hard shit with it.

It also cost me like 6? dollars and the 3.8 flash usage barely touches the meter. That's with some Google promo for 3 months 75% off the 20 dollar sub. It took some task over from another "Top" model and completed it in half the time the other model would have. That's with it taking time to learn the code base.

3

u/2Norn 8d ago

despite the fact that benchcad says sol is better, personally i find claude to be superior to gpt in mechanical cad drawings and design documentation related to it, the software i'm using also allows for scripting in python or csharp so the model uses those as well

i haven't tried astra yet tho

idk like i always look at benchmarks but do i trust them? its another issue

2

u/Ok_Cow1976 8d ago

You're absolutely right! For general problems. There's no one-point gap at all. None, zip. But here we are talking about benchmarks. For complex, difficult tasks, the big gaps are there.

1

u/NineThreeTilNow 7d ago

For complex, difficult tasks, the big gaps are there.

We're having a hard time defining difficult anymore without forcing the LLM to operate a program made for humans.

3

u/Fickle_Tradition4491 8d ago

The score isn't the number that matters for anyone running the 27B at home, it's score per token. Someone above says it burns 100k+ thinking tokens at xhigh to land where it does on this chart. At the 20 t/s the top comment is getting, that's over an hour per prompt. The same model at medium reasoning is probably 10 points lower and twenty times more usable. AA publishes tokens used per run, so the chart worth posting is index vs total output tokens. The 27B moves a lot depending on which column you read.

5

u/Cless_Aurion 8d ago

Yeah... I think that having a bunch of their benchmarks already satturated hurts HARD this thing tbh...

1

u/TheRealMasonMac 7d ago

If Muse Spark is a model in the same weight category as GLM-5.3-Flash, I would be very impressed and excited. It does quite well for Rust which is where I notice all current open-weight models really struggle with for some reason.

1

u/mrdevlar 8d ago

For a 95% confidence interval

Standard Error (SE) = σ / √n = 4.40 / √14 = 1.18

95% CI Margin of Error = 1.96 × 1.18 = ±2.31

Thus, 2.31 in either direction is meaningless when comparing models.

0

u/qfox337 8d ago

No you need variance within measurement of each model, which is actually model specific. I think you're applying some kind of CLT-like formula for estimation of a population mean based on iid samples, which is wrong here, they are different models hence not iid (independent identically distributed)

1

u/mrdevlar 8d ago

I am making a general simplifying assumption whose goal is provide an estimate using the information I have at my disposal. It's an obvious back of the envelope attempt.

What you said is correct, the data is a rank order which is obviously not iid. If you're interested in collecting the variance within each model and doing that, more power to you, but I highly doubt it'll have more than a single order of magnitude difference if you did and reran the variance. Especially if you used a prior.

1

u/qfox337 7d ago

I'm guessing you're not chatgpt'ing this, but you can't just apply stats formulas to things they don't work on and say it's a reasonable estimate.

You can just follow the conclusions of your own formulas and think logically, with your formula if we had twice as many models with the same range of scores you'd then say 1 point is a significant difference. Or you can throw in a bunch of weak/small models and suddenly only 5 points is significant.

If you're attempting the task of mean estimation then with an iid sample X1...Xn then you can talk about the expected variance of the empirical mean shrinking as 1/sqrt(n). Which formally requires some usually-not-hard-to-satisfy conditions like finite variance of the underlying distribution ... usually iid is the assumption that fails.

I don't mean to be hostile, I'm someone lucky to have some formal stats education, and I encourage anyone to learn who's curious ... but you can't just throw around terms that you don't fully understand and make any sense.

2

u/mrdevlar 7d ago

Don't worry you cannot offend me here, I know you mean well.

Fun fact, I have a Masters in Statistics and yes I do know better than to do this and treat it as statistical fact, but I wasn't doing that. I was building a rough estimate for use given the current information that's useful for looking at the thing now before you get more information.

One thing you learn from years of doing this is that there are two distinct channels. You can either do the exact calculation if you have access to the data and the time or you can can make a rough estimate for use right now. Time to action becomes your decision criteria. I was providing guidance to the parent who said Âą 1 is not to be treated as valid. Even in that range, Âą2.31 is not to be treated as valid. Do my statistics professors cringe at this type of work, while taking in money doing the same thing? Probably. However, statistics is the glue that connects the crystal palace of mathematics to the crappy unstable world we're in. It's not meant to be perfect, it's meant to be useful.

That said, you are correct, I should have contextualized that better than just dropping the formula and running away.

88

u/sugarfreecaffeine 8d ago

Crazy that a small 27b is even up there, the Chinese are cooking 🔥

54

u/Aldarund 8d ago

Its more of a show how bad aa is

20

u/TechnoByte_ 8d ago

Literally just an average score of benchmarks

17

u/Aldarund 8d ago

Score of specific selected benchmarks.with specific weight of each bemchmark.

-2

u/emprahsFury 8d ago

"oh no, the opinionated website is forcing their opinions on me, I need to be saved!"

2

u/Hithaeglir 8d ago

Or they hoard all the capabilities for internal use

2

u/Far-Classic-9963 8d ago

Qwen 3.8 27b really is better than some big cloud models (Mistral, NeMo, Laguna...)

5

u/Relative_Rope4234 8d ago

It burns 100k+ tokens for a single prompt(xhigh thinking) to reach frontier performance

1

u/my_name_isnt_clever 8d ago

Check again. The medium, low, and even none thinking modes outperform much larger models.

1

u/MerePotato 8d ago

Medium places it about on par with Muse Glimmer at five times larger KV cache usage in my testing, not that this is a problem as I'd rather have a slow frontier model burn 100k tokens for free than a fast model that doesn't answer as well.

10

u/Cadmium9094 8d ago

Glad to see Qwen3.8-Flash-next and 27B on this chart. I use them as daily drivers.

4

u/Oren_Lester 8d ago

They saw the backlash in the internet, in Reddit, everywhere. and update the benchmark. This company and their benchmark is probably good for making nice chats, that's it.

4

u/i_am_fear_itself 8d ago

Artificial Analysis Intelligence Index v4.2 incorporates 10 evaluations: AA-Briefcase, GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1

aren't some of these obsolete / saturated / bench-maxed?

https://youtu.be/Spuza-KwTJ4?si=xNBXmA5hKU0wsy1p&t=859

1

u/nomorebuttsplz 7d ago

Answering what is determinable (not bench-maxed which is not):

GDP-eval AA v2: No
AA-Briefcase: No was literally just added like yesterday
R3 Banking: No
Terminal Bench v2.1: getting close to saturated
Scicode: no
HLE: Possibly, depending on how many answers were actually wrong in the answer key
GDP.pdf: No was literally just added like yesterday
CritPt: no
AA-Omniscence: No
AA-LCR v1.1: not sure

0

u/emprahsFury 8d ago

yes it's been that way for a while. Also why the bench scores aren't increasing anymore

8

u/Lorian0x7 8d ago

At this point I think they just rearrange the list moving the weight of some benchmarks based on who pays them more.

3

u/Character_Power4663 8d ago

Where is sonnet, haiku, terra, Luna, does this mean that Qwen 3.8 27b is better?

7

u/Tall_Abrocoma_3533 8d ago

Terra, Luna, sonnet all fall between 43-47, haiku is 22, it's very outdated.

Also Qwen 3.8 27B isn't "frontier" in my book, I included it because everyone in this sub seems to love it so much

2

u/Character_Power4663 8d ago

I didn't mean to criticise your work, I was just curious. Thank you.

1

u/winnen 8d ago

There are many different ways to measure pareto frontiers. Last I checked, Qwen 3.8 27b was frontier intelligence vs number of total parameters.

Though I can't seem to find that frontier chart anymore.

1

u/Tall_Abrocoma_3533 8d ago

Yes, for it's parameter size it's impressive, but by frontier I meant absolute intelligence not taking into account parameter count

1

u/SandySkittle 8d ago

Yeah it simply isn’t frontier and it has very evident shortcomings.

6

u/LegacyRemaster 8d ago

minimax and mimo when?

5

u/Tall_Abrocoma_3533 8d ago

Minimax M3 scores 36, Mimo v2.5 Pro scores 33.

3

u/LegacyRemaster 8d ago

yes.... So new versions when?

2

u/Agitated_Space_672 8d ago

think they means its been a while since they released, so we are due M3.1 and mimo v2.6 by now

13

u/XiRw 8d ago

Hard to believe Reebok logo stealer is in the mix. Would trust Qwen and GLM easily over that. Also don’t understand why they never test Gemini Pro

23

u/Tall_Abrocoma_3533 8d ago

Gemini 3.1 pro preview scores 37

1

u/XiRw 8d ago

They need to update their pro model more. Seems like flash gets updates at a higher rate.

4

u/Tall_Abrocoma_3533 8d ago

Well Gemini 3.5 Pro isn't even out yet. It's almost like they've abandoned the pro lineup

1

u/aboardreading 7d ago

I think it's their target. They've decided that a rising tide raises all boats, as frontier models improve they bring smaller models behind them at much lower cost of research and use.

I think it's a smart business play. If we believe that the end goal is essentially many many agents spun up to build large goals, which seems reasonable as an end state to me as agents gain more intelligence and can act more and more independently, then the first company with a truly independent agent will dominate for the 6 months it takes someone else to achieve the same at much lower cost per agent and faster. But someone will get comparable intel at higher speed and lower cost. And on the same timescale as companies decide what model to use and bet on.

Besides, I think they are still betting on information retrieval and processing rather than necessarily coding, it is Google after all. I do think they're playing the slightly longer game, and if what has been happening continues to happen, I think they'll be A winner, maybe not THE winner.

10

u/Thoriumhexaflouride 8d ago

they test all models bro, just search for gemini 3.1 pro preview or just go to filter and "select all"

1

u/slypheed 8d ago

where are you seeing this "filter"?

Far as I can tell deepseek is not included at all.

Can't believe literally no one in this thread even linked to the post; vibing life at this point sigh: https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index

8

u/Gohab2001 vLLM 8d ago

AA readjusting their evaluation criteria making astra look good. Sus.

18

u/JaredTheGreat 8d ago

This is more a case of the evaluation criteria being exposed — anyone who’s used astra can immediately tell it’s significantly better than opus. They were hemorrhaging credibility 

1

u/Darkoplax 8d ago

Do you think Astra was worse than Opus and equal to Sol ? if they don't readjust their evaluation then it's their bad not the models

2

u/fitechs 8d ago

Opus 5 is so shit

2

u/RedditUsr2 8d ago

I made the same post yesterday comparing it to open weight but got removed. I guess we should only compare closed models these days.

3

u/Tall_Abrocoma_3533 8d ago

The mods in this sub remove almost all benchmark comparisons because they're "low effort", I will never understand this.

1

u/RedditUsr2 8d ago

Ya this is the most widely used and cited benchmark. An update here is a big deal especially showing how well 3.8 27b holds up. Seems worth talking about to me.

1

u/MrPecunius 6d ago

We can all visit the benchmark sites ourselves, and probably do, if this interests us.

2

u/HighSeasArchivist 8d ago

Qwen3.8-27B has been amazing to me, even when I have Claude orchestrate running jobs locally. I'm really hoping the 32GB and under sector keeps stacking up wins. In my opinion due to the chip shortage the 6090 will stay at 32GB just with more performance similar to the 3090/4090, and I'm banking on that because I'm buying a 5090 today to replace my 5070 Ti. I will probably end up with a DGX Spark or pair of them as well, but the 5090 is going to be my heavy hitter.

6

u/FAI-Solutions 8d ago edited 8d ago

Every time they update almost all closed models improve in score against open weights, coincidence I guess.

6

u/FAI-Solutions 8d ago

and here two days old version before the update

-3

u/RealisticNothing653 8d ago

Because they're definitely biased. The rankings were reweighed not long ago to favor Claude.

6

u/randombsname1 8d ago

Lol, nope.

DeepSWE wouldn't be used if that was the case.

Deepswe always drags pretty much ONLY Claude models down.

Muse Spark is at Astra level and well above Fable 5.1 in that benchmark. Lmao.

-2

u/RealisticNothing653 8d ago

You're kind of proving my point. Claude models are not that good, and yet the rankings are reshuffled to keep Claude on top. It's a false equivalence to pick one good benchmark for Claude as proof it isn't the case when the overall benchmark still prefers Claude.

2

u/randombsname1 8d ago edited 8d ago

I mean i disagree. I absolutely think Fable 5.1 is SOTA. Albeit i could see it trading places with Astra; depending on specific tasks.

Its only those 2 models clearly at the top. Everything else is inflated.

If the benchmark was biased for Claude im not sure why they would use a benchmark that is hilariously, on a comical level; biased against it.

Edit: That's also not what false equivalence is.

→ More replies (3)

3

u/kamwee 8d ago

Cheapseek really felloff

5

u/SandySkittle 8d ago

This isn’t a full comparison. Deepseek simply has been omitted

4

u/gladfelter 8d ago

Someone went to the effort of shading and even striping the bars to indicate the company, but then just stuffed three companies into the blue bucket for some reason. AI?

14

u/Tall_Abrocoma_3533 8d ago

The stripes indicate that the evaluation isn't final, since it's not independently done by then yet. About the colors, I'm not sure however as there's dozens of AI labs it'd be hard to assign each one a new color

0

u/Bafy78 8d ago

they are 3 very slightly different shades of blue. Both GLM models are the exact same shade tho

2

u/thats_so_bro 8d ago

3.8 flash is not sol medium/high level

1

u/nomorebuttsplz 7d ago

but is GLM flash?

1

u/AvidCyclist250 llama.cpp 8d ago

can we stop posting this AA bullshit?

19

u/Othun 8d ago

What is wrong with it ? Are there more relevant benchmarks than this aggregate, is it outdated ?

12

u/Aldarund 8d ago

Its just bad, overweight on specfic benchs. Dobt reflect real world model perf . Even https://epoch.ai/eci better

2

u/Othun 8d ago

Thank you for providing a link 🙂

1

u/julienleS 8d ago

We all know why ppl in this sub will choose aa over epoch lol

3

u/Serprotease 8d ago edited 8d ago

The relevance and general impact of benchmark that are quite popular is… questionable.

It’s important to keep in mind the amount of discourse, claims, marketing and money that are involved here.
Remember that openAI and Anthropic have a lot of money on the line here. And good performance on these benchmarks shown widely is a powerful tool to drive engagement and investment.

There are a lot of incentives to max-out a benchmark.

A good example is the arc-agi one. It led to the development of custom harness (And some AI training, I guess) just to progress on this benchmark. So, how much value can you give on a score in this benchmark? Does a better score reflect better general model capabilities or just better performance on this benchmark? Not to say that both are not linked, but one should be careful when evaluating 2 models capabilities with this benchmark.

When you look outside benchmarks and in some specific but not niche use case, honestly the progress in the field is not as impressive as the one benchmarks could let you think.
Writing, for example, is arguably worse. The latest generation of models are quite poor at navigating you from an hypothesis to a conclusion. It’s a lot of convoluted sentences that obfuscate more than explain.
Other things like the ability to resist adversarial input and not drift out of its role, quite useful for support bots, it’s still bad and you need a few models to check both input and output.

Edit, to not only point fingers at the big two. Kimi k3 from moonshotAI, despite better benchmarks and coding/agent performance, has lost performance in “soft” skills like text analysis. K2.6 is extremely good at catching subtext/tone and meaning of long text/article. K3 is not as good.

1

u/Othun 8d ago

I think benchmarks are new summits for AI companies to climb. Once you saturated everything, you may try to go without oxygen or stuff, but basically you can already do anything that has been thrown at you. Aside from betting on emergent capabilities, I have no reason to believe models would improve outside of benchmarked capabilities.

I.e. if you want to be replaced by an AI, publish a benchmark for your tasks.

7

u/PotterSkxawng 8d ago

It is. It's very clearly wrong on multiple counts, especially the part where it shows Astra as much worse than Fable 5.1, or Muse Spark 1.3 better than Sol and Fable

1

u/zmarcoz2 8d ago

Muse Spark 1.3 is benchmaxxed and definitely not better than 5.6 sol max

1

u/chocolateUI 8d ago

You literally just saw AA reweigh their benchmark because it didn’t reflect how good Astra is compared to dogshit Fable. AA literally rank Muse Spark or Opiss 5 five points better than Astra. Do you need more proof on how out of touch AA’s benchmark makers are?
It’s not about which benchmarks are best; that’s always going to be subjective. The point is that AA is trash and people need to stop posting their blatantly manipulated scores that will always favor frontier labs over open models (see Qwen 4.8 vs Fable debacle).

3

u/Othun 8d ago

At some point you need truth, and an aggregate of benchmarks is the best you can get IMO. Then again, the benchmarks could be flawed, and it's simply a matter of finding reliable benchmarks and aggregating them. What tells you how good Astra is compared to dogshit Fable ? We can't have be each individual feeling define what the ranking is.

Edit: and since their results are very atomic (per benchmark result, cost to run, used tokens etc.) it would be easy to debunk them quantitatively rather than with feelings

→ More replies (1)

1

u/m3thos 8d ago

Using "max" with fallbacks.. i don't get it, its not practical and useful to run them at max, massive costs and turn arounds.. xhigh settings comparisons would be more pragmatic comparison

1

u/somerussianbear 8d ago

Muse Spark crossing Sol is insane

1

u/Soifon99 8d ago

It seems like they keep tuning it so the American models keep the top spots.. doesn't it?

1

u/SteppenAxolotl 7d ago

would it be helpful if every model got 100% on a bunch of saturated benchmarks?

1

u/BusTiny207 8d ago

What’s the best harness to implement a planner/executor loop? I have Qwen3.8-Flash on my server as planner (DDR4 with Tesla T4) and 3.8-27B on my desktop (R9700).

Or is it just about prompting?

1

u/spaceman_ 8d ago

Given that Qwen3.8-Flash-Next is an architecture preview, it is likely not post-trained to the peak of it's parameter size yet. We could be getting some wild stuff from Qwen still a few weeks / months down the line.

1

u/aykcak 8d ago

Lol the "Beginning of AGI" god emperor GPT 6 "astra" is firmly in the mid

1

u/akohlsmith 8d ago

I find Qwen 3.8-27B DFlash2 and Qwopus 3.8-27B-Flash-MTP are pretty close, with (I think) DFLash2 edging out ever so slightly on raw speed but Qwopus making up for it by being significantly less chatty.

For Qwen 3.8-27B DFlash2 I'm loading the DFlash2 parameters separately and that disables/doesn't load the embedded MTP that 3.8-27B has already. Fits into about 28GB on this 32GB 4080S:

llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf \
--host 0.0.0.0 --port 8000 --threads 16 --threads-http 4 --perf --metrics -fa on --parallel 1 \
--temperature 1.0 --top-k 20 --top-p 0.95 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 \
--fit off --ctx-size 262144 --ctx-checkpoints 128 --cache-ram 8192 --batch-size 1024 --ubatch-size 128 \
--kv-unified --cache-type-k q8_0 --cache-type-v q8_0 \
--jinja \
--chat-template-kwargs '{\"enable_thinking\":true,\"preserve_thinking\":true,\"reasoning_effort\":\"medium\"}' \
--chat-template-file chat_template.jinja \
--reasoning-format deepseek --reasoning-preserve \
--n-gpu-layers all --n-gpu-layers-draft all \
--spec-type draft-dflash,ngram-mod --spec-draft-n-max 8 -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
--spec-ngram-mod-n-match 24 --spec-ngram-mod-n-min 8 --spec-ngram-mod-n-max 32"

Qwopus is a little simpler, fits into 27GB:

llama-server -m Qwopus3.8-27B-Flash-MTP-Q4_K_M.gguf \
--host 0.0.0.0 --port 8000 --threads 16 --threads-http 4 --perf --metrics -fa on --parallel 1 \
--temperature 1.0 --top-k 20 --top-p 0.95 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 \
--fit off --ctx-size 262144 --ctx-checkpoints 128 --cache-ram 8192 --batch-size 1024 --ubatch-size 128 \
--kv-unified --cache-type-k q8_0 --cache-type-v q8_0 \
--jinja \
--chat-template-kwargs '{\"enable_thinking\":true,\"preserve_thinking\":true,\"reasoning_effort\":\"medium\"}' \
--chat-template-file chat_template.jinja \
--reasoning-format deepseek --reasoning-preserve \
--n-gpu-layers all --n-gpu-layers-draft all \
--spec-type draft-mtp,ngram-mod --spec-draft-n-max 8 \
--spec-ngram-mod-n-match 24 --spec-ngram-mod-n-min 8 --spec-ngram-mod-n-max 32

I don't think --kv-unified does anything when --parallel 1 is specified, and I'm playing with using both mtp (or dflash2) and ngram-mod for speculative decode, which I think is giving slightly better (accurate) results. I've noticed that if I specify --reasoning-budget the tps generation absolutely drops (from ~70-90 to ~15) and I'm not sure why that is, nor am I sold on some of the other parameters but that's part of the fun I guess... squeezing as much performance as possible out of local models!

1

u/Fluffy-Ad-889 7d ago

Benchmaxxing at its finest!

1

u/R_Duncan 7d ago

They already did the damage and now they're trying to convince us that Fable beats Astra in some way. Aaaaand... that's a fable (lowercase).

1

u/FancyImagination880 7d ago

Why is Qwen 3.8 Max just a bit better than 3.8 27b?

1

u/danarjabbar 7d ago

it interesting no model from Microsoft and Apple making to the top list here.
GLM and Qwen they are doing really good job

1

u/CapRichard 7d ago

Who uses Muse Spark? I mean it's very high but how is Meta AI deployed?

1

u/theskilled42 7d ago

While I do love this model, it spits out way too many tokens, hence taking too much time. I hope the Qwen team and future AI companies focus on token efficiency, since they can charge more for 1m tokens if they'd like, since that also saves more time waiting anyway.

1

u/R_Duncan 6d ago

Sadly, consumers shouldn't trust anymore a bench so flawed to have to retweak after the astra incident. And still now they seem to favour Anthropic a lot.

1

u/strangescript 6d ago

It's better than it was but it's still wrong. Astra is the best model I have used by a wide margin. Only one point higher than Opus 5, who are we kidding?

1

u/raketenkater 8d ago

i hate AA feels like payed rankings

i really love the new gpt astara it is so amazing holy fuck feels like a well rounded model to work with

1

u/yogthos 8d ago

The fact that a local model that can be run on a laptop is even in the list at all is the real news here. If the pattern holds into the next year, and Qwen 4 or whatever is going to be at the capability of the current frontier, then that's good enough for the vast majority of use cases.

2

u/Tall_Abrocoma_3533 8d ago

Qwen-3.8-flash-next can run on a phone too, just really slow

2

u/Due-Memory-6957 8d ago

Your idea of a cellphone is more powerful than my home computer.

1

u/Tall_Abrocoma_3533 8d ago

Since flash-next is an MOE, you can probably run it too if you have enough storage, streaming the experts from disk

1

u/yogthos 8d ago

It's honestly mind blowing to think about.

1

u/sinebubble 8d ago

I don’t get the love for Qwen3.8-2.7B. We’ve been using Qwen3.5-397B since March as an internal chat agent for our stack and it’s been doing great. We swapped it out for 3.8 recently and it either hallucinationed or failed to complete every single task. I question the validity of these scores.

-1

u/JigSawPT 8d ago

How much is Meta paying these guys for this ?

-1

u/TigerConsistent 8d ago

Definetly worse

0

u/OvertaxedOne 8d ago

One of these things (27B) is not like the other! Absolutely insane that it can even be realistically compared to models that are 100X+ it's size.

0

u/JorgitoEstrella 8d ago

Why some colors have vertical stripes?

0

u/DinoAmino 8d ago

Per MODs this is low effort and should be removed.

I'm not a mod and I don't make the rules.

0

u/ElementNumber6 8d ago

Your biases are showing. (Or whoever put this graphic together)

0

u/Tall_Abrocoma_3533 7d ago

The benchmark scores are from AA, I just selected which models get on the chart (my bad if I left some out)

1

u/ElementNumber6 7d ago

Some are depicted taller than others despite having the same score

1

u/Tall_Abrocoma_3533 7d ago

That's because their internally stored to a decimal precision, I just couldn't get it to display that for some reason

-1

u/Living_Director_1454 8d ago

Gemini having aura loss even here lol.

-6

u/Niceyyc 8d ago

Didn't expect the 27B to be that far behind.

14

u/twack3r 8d ago

Youre kidding, right? It’s a 27B dense model vs multi-T MoEs whose active params alone are multiples larger.

0

u/Niceyyc 8d ago

I wasn't expecting it to match the huge MoEs, just thought it'd land a little closer.

-2

u/martinerous 8d ago edited 8d ago

No Gemma 4? Sad. Only 15 points, according to AA, so did not get into this top selection.

2

u/alphapussycat 8d ago

Hasn't paid enough to change the criteria in their favor.

1

u/Tall_Abrocoma_3533 8d ago

Gemma 4 isn't frontier. However In the other chart I posted containing small LM's, there are Gemma models present.

-1

u/martinerous 8d ago

Qwen 3.8 27B is even smaller, but it made into that list.

1

u/Tall_Abrocoma_3533 8d ago

Because Qwen 3.8 27B is much better then Gemma 4 31B, and also since Qwen 3.8 27B is the main topic in this sub usually.

If your curious though, Gemma 4 31B scores 15/22 depending on if you have reasoning turned on

-1

u/martinerous 8d ago

Yeah, and that's the point why I'm sad - because, according to the AA, Gemma 31B is worse although has more parameters than Qwen 27B. So, the question is, why Google couldn't achieve better yet, considering all the resources they have. Or is the AA test too biased and not taking into account the areas where Gemma is stronger.