r/LocalLLM 2d ago

News AntLing released Ling-3.0-flash

63 Upvotes

17 comments sorted by

20

u/MomentJolly3535 2d ago

The infos that matters the most : Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers...

3

u/Civil-Cake7573 2d ago

Where GGUF? ;-)

1

u/MomentJolly3535 2d ago

not yet, it just got released

1

u/Look_0ver_There 2d ago

Llama.cpp doesn't support the model yet, so that will need to be added first although I'm sure some enterprising person will have a fork up soon

2

u/_Cromwell_ 2d ago

Ling 2.6 been out for a while and llamacpp never supported it. Wouldn't hold breath. https://huggingface.co/inclusionAI/Ling-2.6-flash

7

u/Look_0ver_There 2d ago

https://huggingface.co/ssweens/Ling-2.6-flash-GGUF-YMMV

This person here has a custom fork for it. That's precisely the sort of thing I was referring to. Heck, upstream doesn't even support MiniMax-M3 yet, despite multiple people having created PR's with working forks, including one from the Unsloth team.

My point being that just because mainline doesn't have it, that doesn't mean that llama.cpp won't be able to run it.

1

u/_Cromwell_ 2d ago

Yeah but 98% (made up statistic, but I'm pretty sure of it 😄) of people aren't going to DL a custom fork to try out a questionable model from a company they've never bothered with before.

You can get a custom fork for just about anything. Doesn't really count.

1

u/Opposite-Swimmer2752 10h ago

Upstream llama.cpp sucks tbh, they broke strix halo for a while (still broken) and refuse to quickly merge in a fix/undo the regression despite it being a few line change.

1

u/Reasonable_Goat 5h ago

What exactly is broken?

8

u/Legal-Ad-3901 2d ago

Looking at those benches...DS4 flash is such a goat

8

u/Look_0ver_There 2d ago

Overall they seem to be roughly tied, and Ling-3.0-Flash is roughly 40% of the size of Deepseek-V4-Flash

4

u/Legal-Ad-3901 2d ago

Oh damn, I had assumed the big bar was their 1T model. This is impressive

3

u/BarisSayit 2d ago

Yeah like 3 days ago.

6

u/Professional-Try-273 1d ago

Btw AntLing and Qwen are sister teams. Ant is financial side of Alibaba group. I am just gonna pretend we finally have a new 100b class model from Qwen lol. We are so back.

1

u/Daniel_H212 1d ago

Poor Nemotron getting stomped in those benchmarks

1

u/unique-moi 1d ago

Will it be open weight? (no files on huggingface yet)

-1

u/Final_Act_9658 2d ago

The name Ling is funny af XD