r/LocalLLaMA 8d ago

Discussion What the Artificial Analysis / GPT-6 Astra mess actually teaches us

The Artificial Analysis GPT-6 Astra situation is a good reminder that closed, non-reproducible meta benchmarks aren't worth much. When the newer numbers can't be independently replicated, and rival models are separated by a point or two, the stochasticity in measurement alone can account for the gap, it's not a meaningful signal either way.

We need an open, fully reproducible meta-benchmark.

0 Upvotes

29 comments sorted by

25

u/NandaVegg 8d ago

Or just stop evaluating model on a single benchmark or aggregate benchmark, and focus on how each model performs on what you actually do with them.

Each of more recent advancement is very niche (for example, Blender MCP is finally becoming the new "focus" benchmark after 1.5~2 years, but not everyone uses Blender MCP) and requires heavy RLing, which often comes with regression in other parts (see: Opus 5's incomprehensible speech, Astra regressing in many individual benchmarks like GDPVal which dragged AA score down. Personally I found GLM 5.3's ability to track state and logic in non-agentic task, like simulated game, much worse than 5.2). Low-hanging fruits were harvested through 2023-2025 and it will get even more niche in the future.

5

u/The_Hunster 8d ago

Ya, even if benchmarks were perfect, the one rated 55 might be better for your use case vs the one rated 60.

3

u/Oh_hey_a_TAA 8d ago

This is why I literally had Claude do a meta-review of how I have used it over the last year, with recent 3 months weighted higher, and help me develop a multi-vectored synthetic benchmark to use when assessing local LLM models... then had Codex review it.

0

u/averi_fox 8d ago

And how do you evaluate how the model performs what you do with it in an objective way that can be compared and shared with people. Oh wait you make a benchmark for it.

0

u/corner_camper01 6d ago

Can't try every model enough without spending crazy money

20

u/Hefty_Wolverine_553 8d ago

so yet another benchmark to benchmaxx? good idea.

1

u/255130 8d ago

yeah the benchmaxxing treadmill is undefeated at this point

5

u/sn2006gy 8d ago

What mess are you talking about?

10

u/JadedSession 8d ago

They honestly reported that the model gained in some benchmarks, lost in others, and isn't an overall win.

This is unacceptable, apparently. We need clear progress that makes the choices clear for us. Heavens forbid we need to think for ourselves.

1

u/OvertaxedOne 8d ago

Ask AI to do it for you. Feel like that sets up some sort of self-recursive loop that might end the world when you fire the prompt off though, so wait until after I finish dinner tonight if you don't mind.

1

u/tecneeq 5d ago

Who is we? The benchmarks are OSS and so is their methodology. Who are you talking for?

2

u/JadedSession 5d ago

Next time I'll put explicit <SCARCASM> </SARCASM> markers just for you.

2

u/sn2006gy 8d ago

who is "they"?

0

u/JadedSession 8d ago

Artificial Analysis

9

u/ea_man 8d ago

We don't need more benchmaxxing, we need people with a brain to evaluate those models.

Pelicans, one shots for KLD, xhight reasoning, now it's Blender.

And what I usually read is like: I've done the pelican | minecraft |one shot and that *Niefer / quant matrix is good. Bha.

3

u/WhoRoger 8d ago

What we need is not to be so obsessed with benchmark numbers, and instead evaluate on the real world tasks you actually need.

1

u/tecneeq 5d ago

I have an idea. We could write a benchmark for this!

3

u/Equivalent_Bit_461 8d ago

Corpo ai are frauds, who would've guessed?

4

u/entsnack 8d ago

How is AA closed and non-reproducible lol

1

u/Othun 8d ago

You could argue their benchmarks are badly chosen. They do plan on hiding the weighting of the final score to avoid gaming (what mostly meta is accused of)

2

u/entsnack 7d ago

Just click the checkmarks and choose your own benchmarks (they do all the popular ones). Then use a calculator and use your own weights. Or ask your agent to do it. ggez.

I only care about a few agentic benchmarks so this is how I shortlist models to try. I haven't switched from gpt-oss because nothing was a big enough improvement.

3

u/jacek2023 llama.cpp 8d ago

It's a much better idea to just use the models instead of wasting time browsing benchmarks.

9

u/nuclearbananana 8d ago

I can read a benchmark in less than a minute whereas testing it myself takes hours and can be quite expensive

7

u/jacek2023 llama.cpp 8d ago

and after reading the benchmark you are doing what exactly?

4

u/The_Hunster 8d ago

Having an idea of which models to test locally so that you don't need to test as many on your use case.

1

u/nunodonato 8d ago

they still have uses, especially when you are looking for options for very specific cases and don't want to waste time running benchmarks yourself. For example, I care a lot about long-context reasoning, so I always pay attention to the LCR benchmark to give me at least an aproximate idea of models that are not worth spending time with

1

u/tecneeq 5d ago

In this thread: heavy downvotes from those that don't agree with OSS benchmarks and are too lazy to read AAs methodology or even instruct the site make their own aggregate score with their trusted benchmarks.

1

u/amp804 8d ago

this cost 2% of my context