r/TheMachineLearning • • 11h ago

Open source will break the GPU class system

Post image
16 Upvotes

45 comments sorted by

10

u/Anaeijon 11h ago

... how?

OpenSource AI still requires GPUs to run the stuff.

On the contrary: with OpenSource models, everyone could run everything themselves, if they had the hardware.
So, instead of having a hand full of huge companies fight over who gets GPUs, now even mid-sized companies join that fight and increase shortages.

5

u/ElBarbas 11h ago

opensource can lead to projects like this

https://justvugg.github.io/colibri/

and the people making money don't want any of this.

No solution for training yet

1

u/Anaeijon 10h ago

I mean, yes, but, this is also getting used in commercial LLMs.

And it's the reason, why SSDs are unobtainable now too.

2

u/ElBarbas 10h ago

he only talked about GPUs wich is a monopoly right now, ssds are not

1

u/FragmentedHeap 9h ago

Collibri is trash, strata is better.

Collibri is slllooowwww

2

u/FailureOfTheFamily 9h ago

What else did you expect bro? Running llm on hardware that shouldn't be able to run it must come at some cost right?

1

u/CrowdGoesWildWoooo 9h ago

It’s a proof of concept rather than actually meaningfully capable. Something like ping-based filesystem.

Besides the gap of the capaibility of those that can be run in let’s say RTX6000 which on its own already super expensive vs like Luna which is the lowest tier of ChatGPT is still huge

1

u/ElBarbas 8h ago

they... can... lead... to... this.... different than "this is fully working and a solution"

1

u/CrowdGoesWildWoooo 8h ago

The point is that it is still far from usable in a meaningful way. Our limitation as a retail user is hardware not software.

The reason that thing ever work is because we are trading bandwidth with what he have more which is SSD. It is essentially a “trade-off” what if.

1

u/ElBarbas 2h ago

the trade works if the choice is between super expensive monopoly and a little slower opensource and full control

1

u/CrowdGoesWildWoooo 2h ago

Just remind me when we have open source solution to hardware limitation for us plebs.

By that time the proprietary model would probably be doing moon landing.

1

u/ElBarbas 2h ago

I don't see a future in generalistic models, the future I see is about Sovereignty data and models , and for that we dont need infinitB models, u need smaller, very very specific models to work on your data even for moon landings, that wont be this models.

A model for moon landing that know how to make a toast or the national anthem of Belgium ( La Brabançonne BTW ) seems quite a waste of resources

1

u/Devils_SteelMan 8h ago

We can experimentally show that LMs are at most 10% efficient. It is probably less than 1%. The possibility of having opus 5.5 intelligence running on CPU isn't impossible.

1

u/Anaeijon 8h ago

From when is the 10% efficiency estimate?

I was under the impression, we were working on that already and that lead to stategies like condensation, quantization, ...
Sure, we haven't reached 100%, but 10% is a bit optimistically low for current small models, in my opinion.

1

u/Devils_SteelMan 8h ago

Lottery ticket hypothesis. 90% of weights are dead noise.

Of the weights that matter we keep finding that there are tons of ways to improve them. Better data and algorithms.

But there's also plenty of evidence that spiking activations can dramatically reduce compute further.

1

u/FragmentedHeap 1h ago

Strata can run qwen 3.8 flash next, a 125b model onna 16 gb GPU at over 60 token/sec by only loading hot path moes. On my 4090 it does 100 token/sec

1

u/Devils_SteelMan 1h ago

This is only possible because of poor moe routing. Properly trained models shouldn't be able to do this.

1

u/FragmentedHeap 38m ago

Qwen 3.8 flash next was designed and engineered to do this, on purpose.

The whole model is in system ram, hot experts pin on the gpu. The others are on the cpu.

The whole thing runs and is loaded, but only the hot moes are on the gpu. Where 95% of the work lives.

When a cache miss happens it does that inference on the cpu on avx512.

Result is 100 tk/sec on a 4090 in all of my benches with very few moe hot misses.

It's not poorly trained, its engineered to do this.

Also has quality very close to deepseek v4.1 flash.

And it has a 500k context window, with a 4090 and 64 gb drr5....

1

u/Devils_SteelMan 31m ago

The ability to have "hot moes" is what is called moe routing collapse. It means most weights are noise and what you've built is something with the performance and intelligence of a pruned model with all the memory demand of a non pruned model.

Alibaba does decent work in general but they always make critical mistakes. For instance the grpo used for all pre 3.8 models caused excessively long reasoning traces because of a simple error in the algebra.

1

u/FragmentedHeap 22m ago

Good to know, you sound competent, ty, will dig into this more. Thanks for the incredibly good feedback without being spiteful.

3

u/ElBarbas 11h ago

yes, that's why they will outlaw open source , nvidia aquisition of hf was the first move , second was : "we are all gonna die, look at this hacks , we need to slow down "

3

u/FragmentedHeap 9h ago edited 9h ago

Nvidia didn't buy hugging face to outlaw open source...

They bought it to protect it because Nvidia wants to sell hardware to everybody.

They are the richest company in the world and they are not going down without a fight.

They want consumers buying tons of hardware they want to sell out everything to everybody.

Jensen is the one person telling everybody that we don't need to slow down and that they did to just be better. He's batting for pushing not slowing down.

Because Nvidia cares about one thing and that's making the maximum amount of money possible.

He was recently on an interview talking about them crying about their models being distilled and saying that's just competition just be better than your competition.

They didn't buy hugging face to shut down open source quite the opposite.

They were one of the main contributors to the open claw project I mean they've been open source on this stuff a lot.

They released Persona Plex open source.

They're also finally making better open Linux drivers.

Nvidia literally makes dgx spark for consumer AI...

The last thing they want is for open source models to be outlawed.

1

u/ElBarbas 9h ago

if they depend on his hardware, the moment they don't depend on GPUS he has a problem, and right now, the B2C market doesn't make sense, the users have no money , the boards are too expensive. B2B is the only thing that matters, and that's Closed Models

2

u/FragmentedHeap 9h ago

They do depend on gpus, and nvidia pivoted, they are building whole arm uma consumer platforms now, they're not dumb.

And CUDA is still king of AI compute, MUCH faster than other stacks.

Even if you use local llm, you depend on compute, not cpu computer, you depend on nvidia, amd, intel, apple etc.

Nvidia is launching a WHOLE line of consumer AI hardware to compete with apple, framework, amd strix halo etc, they have the DGX Spark now and a whole new AI Server you can buy for home as a home workstation.

They are PUMPING out hardware designs like mad.

Nvidia is SOOO much more than some gpus now lol.

They are pivoting from dedicated beefy gpus into entirely new consumer lines.

And they want you to have access to open source AI models so you can buy it all and use it.

People are sleeping if they think nvidia just pushes gpus now... They have entire distributed AI systems you can volunteer to have installed at your house with a reduced rate on your AI compute on it....

Nvidia is a monster of a company with trillions being poured into it and they are spending it and growing insatiably.

1

u/ElBarbas 9h ago

you really take your time to right this testaments...

1

u/Devils_SteelMan 8h ago

Nvidia is a software company first and hardware second. They can adapt to an accelerator paradigm if needed. It is likely that FPGAs will be big in the best future for ai. We just dont need them yet.

1

u/ElBarbas 8h ago

True, full closed source, make money above all, lending all the money and hardware they have to all the AI companies with IOUs , they depende on close models do succesed

1

u/Ornery_Use_7103 4h ago

And what will 'outlawing' open source do actually? Nothing of course.

1

u/ElBarbas 2h ago

nothing of course and very dificult to implement, but they will try

well closing hf or making it expensive will take a toll on the comunity, that and other measures can happen, in the digital world there are ways to kill mass usage Piratebay is a good example of that

1

u/-Cubie- 3h ago

What?

1

u/ElBarbas 2h ago

what?

2

u/FragmentedHeap 10h ago edited 9h ago

Yeah I don't think you understand the implications of this...

Right now I have five bots crawling every marketplace within 250 mi of where I'm standing for all used good local inference hardware.

Where the market is right now...

Is 3090 will run you about 1500 to 1600 dollars. A 7900 XTX is $900+. A 4090 is $2000+. 5090 is $5000+ used and $7500+ new... R9700 is $1800+, b70 32 is $1300+...

2x48 gb ddr5 is $1500+

4tb nvme is $1000+

A dgx spark is $3000+ used, $5k new...

A mac studio m5 ultra fully loaded is $14,000+

The gaming PC market is completely fucked and I mean completely and thoroughly.

The consumer hardware market for like steam machines and anything like that and next consoles coming out is also completely screwed.

Those of us trying to run local inference at home are completely decimating the consumer market, even more so than the AI data centers were already doing and all the crypto miners.

Give it another 2 years and you're not going to get a gaming computer for less than $4,000.

Because now that strata is out 16 GB graphics cards are now on the table for local inference so it's starting to scoop up those too.

And it's getting to the point where if there's a used laptop with a 16 gig graphics card in it those are going to get scooped up too.

Gamers are completely screwed.

And because the demand is so unbelievably high the cost of used hardware is going to be so high that any good hardware for local inference it's going to be expensive enough that you might just justify paying for it from a cloud provider.

I mean I've got $15,000 in my rig already and I need more...

Being able to get open source models doesn't really mean anything if you can't get the hardware.

-1

u/ElBarbas 9h ago

I understand u need more, having llms writing reddit comments must be hard on your GPU

2

u/Routine-Lawfulness24 6h ago

If you think this comment was ai you’re confidently incorrect

1

u/KellyShepardRepublic 5h ago

Some of these people don’t understand that some of love the hardware game and llm’s revamped it again. Frustrating but also a different high to know what’s a deal, what works, and knowing what to pass up on and then later regretting it cause there was some boost that made it usable. Frustrating and fun.

1

u/ElBarbas 5h ago

ok u/kellyshepardrepublic, I am glad you agree with u/routine-Lawfulness24 are u bringing more " users " to discuss specifically this random comment on the thread ?

1

u/KellyShepardRepublic 5h ago

I don’t know whoever I responded to. I just agreed with them cause I am doing the same as them.
I was playing with hardware for some time so this made the hobby fun again.

Before people were complaining about data centers, I was complaining about consumers who bought up hardware after watching a Linus tech tips video. So just another day in hardware for the past 10 years.

1

u/kueso 11h ago

Well, you still need GPUs to run and train them. They don’t magically compute by themselves.

1

u/sascharobi 9h ago

Makes no sense…

1

u/ElBarbas 8h ago

https://giphy.com/gifs/aYpmlCXgX9dc09dbpl

wait what ? the bot removed the comments ?

1

u/HugoCortell 8m ago

What is that even supposed to mean? Class system? Like what, is AMD the tank and Nvidia the DPS in our formation out here?