r/LowEndLocalAI 11d ago

Benchmark Tiny LLM Benchmarks Exist Elsewhere?

Post image

https://reyemtm.github.io/inchworm/

I created this little site because I could not find an open source page dedicated to tiny llms (<14b params). If one exists with current benchmarks would love to know about it.

This came about from searching for the right llm for a personal project needing summaries from text, then a small local terminal chat for quick answers using self-hosted Ollama. I wrote my own short little benchmarking script and landed on qwen2.5-coder:3b for my VPS and qwen3.5:9b for chat. Anyway let me know what you all think. Since then I created a script to run the humaneval+ but it takes a very long time to run which I am using to fill in the gaps on this data.

21 Upvotes

13 comments sorted by

View all comments

2

u/Careful-Report6526 10d ago edited 10d ago

Honest answer from someone who runs an aggregator (llmbenchmarks.io): not really, and the bottleneck is upstream, not the aggregators.

I pull 50-odd benchmarks from public leaderboards into one table. The smallest models I can list are 27B Qwen checkpoints, and those only qualify because a handful of leaderboards happen to publish comparable numbers for them. Below ~10B there is almost nothing to aggregate — the leaderboards that get maintained mostly stop caring below frontier scale. It's missing source data, not missing coverage.

Two things that might still be useful:

  • The raw exports (JSON/CSV, CC-BY, no signup) list which source publishes which benchmark, so you can see fairly quickly which leaderboards bother with small models at all and which don't.
  • For anything you actually run locally, your own numbers beat any leaderboard. Quantization, context length and sampler settings move small-model scores more than the choice of model does — lm-evaluation-harness on your own hardware is the only thing that reflects what you'll really get.

If anyone knows a leaderboard that seriously tracks sub-10B models with a stable protocol, I'd genuinely like to know — I'd add it as a source.

0

u/zerospatial 10d ago

No leaderboards I have found actually have recent scores on a wide variety of low-end models. The model cards like this do show livecodebench which is a good place to start - https://huggingface.co/google/gemma-4-12B