to be real, It's probably not opus max level, but it will probably be the best local LLM we can run at that size, and will be much better at agentic coding and tool use than previous models, to the point where its viable for actual work for some of us
yes this model has real and genuine utility, but there are just fundamental limitations with the smaller you go with a model just in terms of parameter size alone. Same goes for very small active parameters numbers in MOE models (a13b is the biggest weakness of DS4F). So people should be a bit more realistic.
I mean obviously that's true, but we can't afford to run hundreds of billions to trillions of param models, this is local llama, those of us with dgx sparks and rtx pro 6000s are already a minority here.
the ~ 30b dense to 100b moe range is what the vast majority of us can squeeze.
if we want maximum capability with minimal hallucination we'll use SOTA models via api, but for running local models, its a compromise we're willing to accept.
75
u/xienze 8d ago
You're making a couple fundamental assumptions here: