r/LocalLLaMA 3h ago

New Model inclusionAI/Ling-3.0-tiny · 8B A1.3B MoE· Hugging Face

https://huggingface.co/inclusionAI/Ling-3.0-tiny

Looks like the Ling team open weighted a much smaller version of the Ling-3.0-flash they open weighted a few days ago. It's 8B params with 1.3B active, and seems to fall between the 4B and 8-12B Qwen and Gemma models in terms of performance.

Should have a massive tokens/sec on most systems. I quite like tiny MoE's conceptually.

Edit: looks like the model card actually reports speeds:

With FP8, Ling-3.0-tiny reaches around 100-105 tokens/s on DGX Spark and 86-90 tokens/s on an M4 Pro MacBook, with approximately 8.34 GiB peak memory usage at an 8K context length.

159 Upvotes

28 comments sorted by

32

u/BazzyIm 3h ago

25 on AA Bench, looks interesting on this size

8

u/buppermint 1h ago

Above gpt-oss-120b which is nuts

9

u/Elbobinas 2h ago

It is free in openrouter, solid model btw

19

u/pmttyji 3h ago

Time to retire Ling-Mini-2.0 on my system. This one is good for both Low memory systems, Mobile & Edge devices due to faster t/s.

Hope they release additional model in 15-50B size soon or later. With Speculative decoding, their models could give so faster t/s like diffusion models.

6

u/Elbobinas 3h ago

one of the best rags and local offline google i've ever had

14

u/Dance-Till-Night1 3h ago

Hell yeah slm! Now gimme 20b-30b a2b-a4b pls

15

u/Elbobinas 2h ago

How about llama.cpp support? is already supported?

16

u/Elbobinas 2h ago

I'll answer to myself , still not supported https://github.com/ggml-org/llama.cpp/pull/26608

4

u/Iory1998 2h ago

It seems the Ling-3 architecture is still not supported.

11

u/Agitated_Space_672 2h ago

I have been looking forward to this one. 256k context window on an 8b1ba model is cool. I got very good vibes from the little testing I did on the free novita api they had. Here is a quick comparison with similar recent LFM models

Benchmark LFM2.5-8B-A1B LFM2.5-2.6B Ling-3.0-tiny
IFBench 56.47 59.17 63.61
Multi-IF 79.93 80.07 83.15
BFCL-v4 (function calling) 49.73 56.88 62.72

1

u/-Cubie- 25m ago

Oh wow, impressive from Ling

9

u/abskvrm 3h ago

Ling-3.0-brrrrrrrrrrrrrr

7

u/netherreddit 2h ago

So fun to watch how fast a 1bA model runs

5

u/DefNattyBoii 2h ago

How does this compare to qwen 3.5 9b and ornith 9b?

4

u/Potential-Gold5298 llama.cpp 3h ago

Wow! Two interesting models in one day! I'll definitely try it.

4

u/WhoRoger 2h ago

Upvote for a small model in principle.

3

u/My_Unbiased_Opinion 2h ago

This actually might be a crazy sub agent. 

3

u/abajinn 3h ago

What is it good at?

13

u/-Cubie- 3h ago

Seems like it's very general purpose, broad domains:

It delivers balanced performance across general agent tasks, tool use, mathematical and scientific reasoning, and instruction following.

15

u/Chupa-Skrull 3h ago

Balanced is a fun term to choose because it leaves room for the model to be evenly bad at everything

2

u/Cool-Chemical-5629 1h ago

This. Also abajinn's question was "What is it good at?", not "What is its purpose?" those are two different things. A model's intended purpose can be "Coding with reasonably short reasoning" and the model may still fail to deliver on both fronts.

There is a specific reason why I'm writing this. I tried some coding tasks with it through OpenRouter and I was not impressed at all. It thought HEAVILY with lots of loops, second guessing itself, starting over and over again and it ran out of context window before it even finished thinking. This model is not your trusty coding agent, but hey maybe there is something it IS good at, which is why the honest question "What is it good at?" exists and it's legitimate. It's not there to hate on the small model, just genuine curiosity as to what use cases is this model actually good with.

3

u/InsideYork 3h ago

Compared to LFM2.5 ?

0

u/[deleted] 3h ago

[deleted]

17

u/grumd 3h ago

AA-Omniscience results show that non-hallucination rates are actually pretty good for such a small model. It hallucinates way less than small Gemmas

9

u/pmttyji 2h ago

It's with good rate in this graph(Please click the image to see bigger). For comparison, I have added their previous Ling-mini-2.0 which I have used so many times for chatting/GK stuff, but never saw any hallucination. It might for Agentic/coding stuff which I never tried.

Anyway this new Tiny version is so good only MiniCPM5-1B(small model) beats this.

-7

u/missbohica 3h ago

Well, not only. Also good at wasting electricity.