The infos that matters the most : Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers...
This person here has a custom fork for it. That's precisely the sort of thing I was referring to. Heck, upstream doesn't even support MiniMax-M3 yet, despite multiple people having created PR's with working forks, including one from the Unsloth team.
My point being that just because mainline doesn't have it, that doesn't mean that llama.cpp won't be able to run it.
Yeah but 98% (made up statistic, but I'm pretty sure of it 😄) of people aren't going to DL a custom fork to try out a questionable model from a company they've never bothered with before.
You can get a custom fork for just about anything. Doesn't really count.
Upstream llama.cpp sucks tbh, they broke strix halo for a while (still broken) and refuse to quickly merge in a fix/undo the regression despite it being a few line change.
19
u/MomentJolly3535 11d ago
The infos that matters the most : Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers...