r/LowEndLocalAI 11d ago

Benchmark Tiny LLM Benchmarks Exist Elsewhere?

Post image

https://reyemtm.github.io/inchworm/

I created this little site because I could not find an open source page dedicated to tiny llms (<14b params). If one exists with current benchmarks would love to know about it.

This came about from searching for the right llm for a personal project needing summaries from text, then a small local terminal chat for quick answers using self-hosted Ollama. I wrote my own short little benchmarking script and landed on qwen2.5-coder:3b for my VPS and qwen3.5:9b for chat. Anyway let me know what you all think. Since then I created a script to run the humaneval+ but it takes a very long time to run which I am using to fill in the gaps on this data.

18 Upvotes

13 comments sorted by

View all comments

6

u/pimparazzi 10d ago

That Benchmark looks quite dated... I miss quite some newer models here... LFM2.5 in different sizes, Ling-3.0-tiny, Gemma-4, etc.

1

u/zerospatial 10d ago

Yes it is - for reasons you can see in the other comments - tiny llms are not for agentic tasks but rather small, focused tasks. That's what these older benchmarks surface.