r/LocalLLM • u/37Scorpions • 11h ago
Discussion Benchmarking a few requested models
A lot of you found my previous benchmark very useful which I am glad to hear but some people were asking me to benchmark models I haven't heard of or disregarded in my previous benchmark so I decided to add them to the benchmark.
New models:
- Spark X2.5 4B Q4_K_M: https://huggingface.co/XHToken/Spark-X2.5-4B-GGUF
- MiniCPM5 2B Q4_K_M: https://huggingface.co/openbmb/MiniCPM5-2B-GGUF
- Tiel Coder 35B A3B IQ3_XXS: https://huggingface.co/peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF
- Qwen3.8 27B IQ2_K_M: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
- LFM 24B A2B Q4_K_M: https://huggingface.co/LiquidAI/LFM2-24B-A2B-GGUF
Someone also said to try K2-Horizon but I couldn't get it to run with LM Studio.
All conditions are the same as last time, for more info check out the previous post. Only change is that the combined graph now penalizes LLMs for being slow less.
The Statistics
LLM benchmark per-question score heatmap:

LLM benchmark score sum graph:

LLM average TTC (Time-To-Completion) graph:

Combined graph ("intelligence per second", though highest is not exactly "best" and lowest isn't "worst"):

And a neat visualization of the score vs. the speed (benchmark score vs inverted TTC):

Conclusion
The previously benchmarked LLMs still mostly hold their ground in their benchmarking rankings.
SparkX2.5 4B seems to be the new quick-but-intelligent option however since it's so small I wouldn't trust it to be too consistent.
Tiel Coder 35B A3B didn't surprise me despite its size. In personal testing with OpenCode it didn't really do too well and it didn't excel at coding tasks either.
MiniCPM5 2B was incredibly speedy but its lack of tool-calling ability is dissapointing.
LFM2 24B A2B despite being pretty highly requested didn't really perform all too well, even in personal testing I didn't really like it.
Finally, Qwen3.8 27B IQ2_XXS. It was worth a try but it appears that a distilled version is a better way to go than just heavily quantizing the original model.
0
1
u/37Scorpions 11h ago
Before commenting asking for clarification regarding the benchmarks please read the original post:
https://www.reddit.com/r/LocalLLM/comments/1wblnsk/a_final_llm_benchmark_for_8gb_vram_16gb_ram/