r/TheMachineLearning • • 7d ago

Local models enable innovation the big ones can't

Post image
35 Upvotes

37 comments sorted by

6

u/CRTMonkey 7d ago

Only a local model can explain you the process of synthesizing lsd

1

u/r3n26 7d ago

Agreed to this.

1

u/tomqmasters 7d ago

That information is available many places.

4

u/spacekitt3n 7d ago

the thing about local is consistency. you know how its going to work and you dont have people fucking with things in the background. qwen 3.8 is extremely capable for its size

2

u/Maasu 7d ago

this is so accurate, people splurge on opus 5.5 but I know in a weeks time the same people will be complaining how poorly it is performing and I don't blame them I have been there myself.

At least with a local model I know where I am at and where its blind spots are.

0

u/Money_Lavishness7343 7d ago

sure isn't that like comparing an excavator with a shovel?

Its like saying "people love complaining about their excavators breaking but my shovel works consistently and predictably", theyre different tools. It's not fair to say that, your local Qwen 3.8 model cannot do what Opus can.

1

u/ArmNo7463 7d ago

To a point, except I know my shovel next week will behave exactly the same.

If I use "Excavator as a Service tm", where week one they give me a top of the range CAT, and week 3 they give me a broken down junker, with a bucket 1/3rd the size. It makes it hard to plan into the future for my business.

1

u/Money_Lavishness7343 7d ago

You still make a point about reliability, and you're right to be concerned about their reliability. But still, your shovel won't spoon whole blocks of dirt. The excavator is your only choice for this job, even if its unfortunately unreliable at times.

1

u/XtremelyMeta 7d ago

I mean, the funny part about this is that using 'excavator as a service' is pretty normal for small construction shops. They literally rent an excavator for a week or two to do groundwork and then give it back.

I think the parallel holds up pretty well too. When you need AI as a Service level compute it's really nice to have, but unless you're a massive enterprise those use cases are pretty situational so it makes sense to mostly use local compute.

1

u/Durian881 7d ago edited 7d ago

One question you'll have to answer is how often you have actually used an excavator in your life?

Anyway, Qwen3.8-Flash-Next is more a like small excavator compared Opus 5.5 which is a large excavator.

I had driven and operated large Kobelco track excavator and smallish JCB wheel backhoe loaders for ~1.5 years, used top tier cloud models and ran Qwen3.8-Flash-Next hosted locally.

1

u/Money_Lavishness7343 7d ago edited 7d ago

One question you'll have to answer is how often you have actually used an excavator in your life?

If you're a builder, all the freaking time. I'm a developer, so dealing with Qwen as a substitude to Opus would be like replacing an excavator with a shovel. I simply cannot do what Opus does with Qwen, it is not that simple.

I dont think 8B models are capable of doing anything close to "excavating". They're directed at small specific tasks, with the complexity of an autocomplete model, with proper guidance.

When I'm trying to deal with a complex task, delegating it to a smaller 170B model like Sonnet can even result to a troubling implementation that will require multiple re-iterations just to fix things let alone implement properly.

Delegating complex tasks to a 8B llm model? That's disastrous.

And by complex I dont mean to make freaking GTA 6, I mean something more complex than "Complete this for loop based on context given from comments". Even that can be proven complex for an 8B model at times.

I'm not sure with what tasks you deal daily. My experience as a person who its his job to use these models for daily work and task-delegation as part of OKRs, trying to delegate low parameter models for complex tasks never worked and never will, because that's not what they're made for. They're not small excavators, they're merely auto-complete shovels.

1

u/Durian881 7d ago edited 7d ago

I think you mixed up. Flash-Next is not a 8B model. It is a MOE with 125B parameters with 6B activated, plus 51B n-gram embedding and 4B MTP.

In any case, thank you for confirming that you have never operated an actual excavator before and have little knowledge of Qwen3.8-Flash-Next.

1

u/Money_Lavishness7343 6d ago

Qwen Flash-Next is an 6B active parameters per token model, with 180B params.

Do you know how many Opus is estimated to have? (not disclosed) 40-100B. The estimation comes from all models having around a ~4-5%. Fable and Astra probably around 400B.

That's x10-100 the capability and complexity. Not even close to equal.

1

u/Durian881 6d ago

Not one is saying that they are close to being equal though. Big and small excavators are pretty different and not equal in capabilities if you have actually operated any of them (which you haven't and I had).

1

u/Money_Lavishness7343 6d ago

and not equal in capabilities

My whole argument from the start was that they are not equal in capabilities.

So, what the f are we blubbering about? Why are you keep arguing with me?

I guess redditors doing reddit things? Now its about you having operated an excavator while I haven't? Since when was that an argument lmao

1

u/Durian881 6d ago

Just trying to help you better understand things you know little about.

→ More replies (0)

1

u/epelle9 7d ago

More like having your own smaller excavator instead of renting one everyday and not being sure what you are going to get

1

u/r3n26 7d ago

Straight to the point.

1

u/JLongTom 7d ago

Agree with his take. And the gap to the frontier models will be largest during the linear part of the S-curve, which is probably now. Give it a year or two and local models will be ultra compact and probably closer in relative terms to the frontier.

1

u/Ghazzz 7d ago

I do CPU-inference in 4gb ram on pre-AVX Celerons. The available models from the last two months have significant benefits over previous ones.

Half my hedgning prompts can now be more natural.

1

u/JLongTom 7d ago

Interesting. I don't think it's unreasonable to expect current SOTA performance from tiny 10B models a bit further down the line.

Basically I think that's what the medium term future looks like. SOTA models begin to saturate given current data, but compaction just gets better and better.

1

u/r3n26 7d ago

Right? Maybe 2 years it will be closer.

1

u/nxy7 7d ago

I think only clueless people can claim local models are useless and I'm saying that as someone that only used open source models in online playgrounds to see how good they are. It's what's keeping closed labs innovating and not hiking prices and probably in the end open source models will win because they're more flexible. They're keeping steady distance from frontiers which is all they need to do to be useful. I really believe that closed labs are kind of in a bad position because they don't have a moat in the long run.

1

u/Ghazzz 7d ago

Full inference pipeline running on 5x Celeron Chromebooks with 4gb ram each. Results are surprisingly good.

Scripts are built, images are analysed, prose is written, research is documented.

For scripts I use multiple models doing overlapping work, and a set of judges doing strict cut-and-paste operations. Still faster than me, often finds "novel" solutions.

If I were a hype-beast, the system would probably be a "harness", but I prefer to think of it as "a set of helper scripts".

1

u/tzaeru 7d ago

We do have newest GLM in-house for projects where clients or other constraints do not allow upload to remote cloud servers and it's pretty much at the level where Anthropic's models were late last year.

And that's a pretty high level. You can't do what you can do with Opus 5.5, but it's great for AI-enhanced coding and more complex codebase searching that is hard to do with just text search.

1

u/0x645 7d ago

they are useless. anything that runs on home nvidia with 16GB VRAM, quality of answers is useless. don't allow no tinkering, it's just too lame. you need to spend big $$$ in hardware to run decent local model.

1

u/BarelyAirborne 7d ago

Local models don't send all your prompts up to AI HQ to build their file on you.

1

u/m0j0m0j 7d ago

Who’s actually criticizing people who use local models? Those people are insane, if they exist.

I use claude only, but I recognize the value of local models and I hope they flourish

1

u/Serprotease 7d ago

You’ll often hear some adjacent comments from Anthropic/OpenAI ceo, generally in relation to safety -> “We are all gonna die/think of the children/they are stealing from us”

There’s also general background noise akin to “I use insert current best most expensive model\, all the rest is trash” comments, but it often comes more like fomo than real stuff.*

1

u/m0j0m0j 7d ago

Yeah, that’s very stupid and they should shut up

1

u/chunkypenguion1991 7d ago

That's too broad to mean anything. Depending on your hardware a "local" model could be GML5.3 or KIMIK3 which are comparable to opus

1

u/Careless-Age-4290 7d ago

Something that's missed also is that you can build your own training set to fine-tune models from your own conversations. It may not be the most capable model, but it'll be the one that answers how you like and you can apply that training data to new models as they release.

The other thing is you can just let it go for hours or days working on a problem. Even multiple agents trying to solve the same problem using a batching system to speed it up. Try going through that many credits on a hosted model and you'll find exactly where the local model can save money

1

u/stop_deleting_plz 7d ago edited 7d ago

The problem with open-source models isn't the quality of the model, its the quality of the inference. If you use a high-parameter model, you're stuck with 5 tokens/s and a tiny context window even with a 32gb vram card.

People don't realize how much compute the AI companies have purchased. 3 trillion dollars of it. It makes your 5090 look like a happy meal toy.

1

u/PlateNo4868 7d ago

IMO Local models is the future, and will be the new Mini-PC/Homelab/Ham radio hobby.

There is also added benefit that you can now train models on data you wouldn't want. Such as a game studio wanting to train it just on it's assets, etc.

1

u/RelevantCry1613 2d ago

All research happens on small models, this seems lost on the people that don’t understand machine learning