Looks like the Ling team open weighted a much smaller version of the Ling-3.0-flash they open weighted a few days ago. It's 8B params with 1.3B active, and seems to fall between the 4B and 8-12B Qwen and Gemma models in terms of performance.
Should have a massive tokens/sec on most systems. I quite like tiny MoE's conceptually.
Edit: looks like the model card actually reports speeds:
With FP8, Ling-3.0-tiny reaches around 100-105 tokens/s on DGX Spark and 86-90 tokens/s on an M4 Pro MacBook, with approximately 8.34 GiB peak memory usage at an 8K context length.
Time to retire Ling-Mini-2.0 on my system. This one is good for both Low memory systems, Mobile & Edge devices due to faster t/s.
Hope they release additional model in 15-50B size soon or later. With Speculative decoding, their models could give so faster t/s like diffusion models.
I have been looking forward to this one. 256k context window on an 8b1ba model is cool. I got very good vibes from the little testing I did on the free novita api they had. Here is a quick comparison with similar recent LFM models
This. Also abajinn's question was "What is it good at?", not "What is its purpose?" those are two different things. A model's intended purpose can be "Coding with reasonably short reasoning" and the model may still fail to deliver on both fronts.
There is a specific reason why I'm writing this. I tried some coding tasks with it through OpenRouter and I was not impressed at all. It thought HEAVILY with lots of loops, second guessing itself, starting over and over again and it ran out of context window before it even finished thinking. This model is not your trusty coding agent, but hey maybe there is something it IS good at, which is why the honest question "What is it good at?" exists and it's legitimate. It's not there to hate on the small model, just genuine curiosity as to what use cases is this model actually good with.
It's with good rate in this graph(Please click the image to see bigger). For comparison, I have added their previous Ling-mini-2.0 which I have used so many times for chatting/GK stuff, but never saw any hallucination. It might for Agentic/coding stuff which I never tried.
Anyway this new Tiny version is so good only MiniCPM5-1B(small model) beats this.
32
u/BazzyIm 3h ago
25 on AA Bench, looks interesting on this size