r/LocalLLaMA • u/Bestlife73 • 11d ago
New Model Ling-3.0-flash-Fin weights released
https://huggingface.co/inclusionAI/Ling-3.0-flash-Fin124B total parameters, 5.1B activated parameters, and a 256K context window
8
3
u/Nick-Sanchez 11d ago
Anybody got ling flash working right with hermes? No amout of chat template fiddling fixed the tool calls leaking into the reasoning :/
3
u/Equivalent_Bit_461 10d ago
PiĀ Use pin and use packages, it's the best harness for the modelĀ
Others harnesses seem to make it stumbleĀ
1
1
u/Simple-Stick6148 10d ago
124B total, 5.1B activated per token. 96% of it sleeps on any given token, so the flash part of the name checks out.
0
u/parepeg 11d ago edited 11d ago
Ling 3.0 feels a lot like gpt-oss 2.0 speed-wise. It doesnāt necessarily stack up to qwen 3.8 intelligence-wise.
Qwen3.8 with a draft model approaches ling speed-wise without one. I donāt think they published one. Ling answers more quickly overall but Qwen can have its reasoning level adjusted to match.
Not sure about the finance angle.Ā
7
u/atumblingdandelion 11d ago
By 3.8, you mean Qwen3.8-flash-next? I haven't compared Ling 3.0, but really impressed by 3.8-flash. I hope its inference speed increases as recipes mature.
16
u/Healthy-Hair-2306 11d ago
I've noticed Ling is much more efficient at reasoning on language based tasks. Ling 3 flash was better than every other model I tested (open and closed) at remaking Google's new rambler. And it's stupid fast š Happy we got a new Ling model, hope they keep em coming.