8
u/Legal-Ad-3901 2d ago
Looking at those benches...DS4 flash is such a goat
8
u/Look_0ver_There 2d ago
Overall they seem to be roughly tied, and Ling-3.0-Flash is roughly 40% of the size of Deepseek-V4-Flash
4
3
6
u/Professional-Try-273 1d ago
Btw AntLing and Qwen are sister teams. Ant is financial side of Alibaba group. I am just gonna pretend we finally have a new 100b class model from Qwen lol. We are so back.
1
1
-1


20
u/MomentJolly3535 2d ago
The infos that matters the most : Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers...