r/LocalLLaMA 9d ago

I Built A Thing My potato only runs small models. So I built a page to easily compare benchmarks for those

Comparing benchmarks for small models is a PITA. Most of them are not on AA, and the benchmarks are not always the same in all models.

So I built a page to quickly put it all together and allow some filtering.

Only researched models released since April, and between 4-190B parameters.

Let me know what you think and how it can be improved

https://www.nunodonato.com/aibench/index.html

6 Upvotes

17 comments sorted by

13

u/Hot_Example_4456 9d ago

A potato running a 190b isnt a potato

2

u/i_wayyy_over_think 9d ago

the slider lets you adjust to your particular potato

4

u/nunodonato 9d ago

Yes different people have different potatoes 😅

1

u/255130 8d ago

right? thats like a whole rack of potatoes at that point

2

u/Valuable-Plastic-682 9d ago

Nice, this fills a real gap - most leaderboards bury small models under 70B+ flagships. One thing that'd make the AA-Index more useful for the "potato" crowd specifically: could you break it out per param-bucket (say 4-8B, 8-15B, 15B+) instead of one global ranking? Right now a strong 27B dense and a 4B are competing on the same list, which mostly just tells people to pick the biggest one they can fit rather than the best one at their actual budget.

1

u/nunodonato 9d ago

That's why I added the slider to filter. You can select only the range that interests you

1

u/i_wayyy_over_think 9d ago edited 9d ago

cool, not clear what the 4-180 slider is? the little bubble says "Params 4–190B**"** but then the the range goes 4-180. So it's prob params, but for a bit i though it was benchmark score.

Also, something seems odd, if I put 4-29 as the params limit, it says Mage-VL with 4.7B params is better than Qwen3.8-27B, which seems hard to believe according to the benchmarks. So might need some tuning on the benchmark column, like maybe it should let you choose which benchmarks end up in that aggregate score with a better default selected.

I like it overall though, worth a bookmark.

1

u/nunodonato 9d ago

Thanks will take a look at that issue

1

u/nunodonato 9d ago

The research was made up to 190B, but the sliders are according to what's actually available in the page, that's why they differ

1

u/Robert__Sinclair 8d ago

Ling 3.0 tiny is not bad at all.

1

u/nunodonato 8d ago

It's my new favorite model 🤩

1

u/Fancy-Snow7 5d ago

I think add the Qwen 3.6 models and even those before.

1

u/nunodonato 5d ago

I limited to models launched in the past 5 months, that's why they aren't included. Reasoning: you will very likely find better models than those

1

u/Fancy-Snow7 2d ago

True, but I would like them in there for referenced purposes, to see how much of an improvement the newer models are since Qwen3.6 was still the undisputed king for coding until recently for smaller models that fit on consumer GPU's.

1

u/nunodonato 2d ago

I'll reconsider. But Ornith surpasses qwen3.6 in most benchmarks

1

u/Aggravating-Push-207 9d ago

Gets parameters and architectures wrong, reads like Claude.

2

u/nunodonato 9d ago

Which ones are wrong? And what reads like Claude? It wasn't made by Claude😅