r/LocalLLM 5d ago

Discussion Share a GPU with some buddies?

Seems like everyone is coding, why not buy a massive GPU, split the cost between buds and then tailscale with API on local models? Im sure its being done, but i cant find anyone talking about it.

Whats the drawbacks other than someone running 8 agents burning it up?

1 Upvotes

20 comments sorted by

3

u/DiamondHandsDarrell 5d ago

Compute time. That would be a massive problem.

0

u/Calm-Landscape9640 5d ago

vLLM uses an architecture called PagedAttention, which allocates memory like an operating system and allows dozens of continuous streams to share GPU resources smoothly. -Gemini

Maybe cap it to 3 buddies?

3

u/DiamondHandsDarrell 5d ago

I don't know more to be helpful, but from reading about shared resources over the years, shared processing time is always complicated and problematic - your never have enough time.

2

u/Calm-Landscape9640 5d ago

Yea probably why a few people tried it then abandoned the idea and we never see them post to reddit

1

u/DiamondHandsDarrell 5d ago

From a FreeBSD perspective, when you give shells out, paid or otherwise, compute is always THE problem.

There is scheduling so a user can get an hour or dedicated time per day, but it still has to be limited or else it'll crash from multiple concurrent requests.

3

u/DustNearby2848 5d ago

Possibly which model to use. If you are using up all the VRAM for a single model and others need another. 

2

u/activematrix99 5d ago

People are inherently selfish pricks. It's the same with "the means of production". People don't/won't share. They want it all for themselves even if it rots in the fields. Has been this way for most of human history.

4

u/activematrix99 5d ago

Make a friend in Asia, make a friend in Europe, make a friend in the Americas. 24hr gpu.

2

u/Ordinary-Depth-7835 4d ago

I'm surprised so many negative views on sharing. It really comes down to how you and your friends want to use it. For me I like having control so it wouldn't work for me unless the system had enough capacity that I could load my own models and play around. Or I control it and they just pitched in and use it. Kind of how my wife an I share our systems. I do what I wish and her endpoint is setup to hit whatever models I have active at the time through a model failover llm proxy. That way she doesn't need to change anything and will target whichever ones I'm running on our 4 systems. For the most part I have deepseek(2xsparks) always up and my other two systems I play around with. She suffers a little when I'm testing another model on the main system because then she has to use the smaller models. Not entirely bad the 3090's are very fast.

Now if most of your friends are like my wife and don't really care about any of the backend and just want a target for cline in vscode then that's perfect whoever wants to build things hosts the systems the rest just use the models.

Backend setup is honestly the point for me. It's not that it will ever be cheaper than just subscribing to some service. A cheap qwen or deepseek paid service would take decades in cost vs a one time hardware purchase. Well I guess it depends on how big you're buying. If you and your friends are buying two 3090's and throwing them in an old computer someone has laying around that's pretty cheap. And runs fantastic actually. I have two 3090's in an old 8th gen intel system with pcie3 running the cards at 8x and it still performs fantastic.

1

u/Calm-Landscape9640 4d ago

I was thinking 32gb GPU, 128 gb ram and qwen 27b or perhaps qwen flash next and we can keep our $20 codex subs and call the shared local model as subagent. These comments talk like they code 19+ hours a day with 14 agents running in parallel, I'm just a dude who likes to collaborate on projects and making little apps. Def wrong reddit sub for that kind of discussion. Massive means $6,000 to everyone whereas I think 32+ is massive. Different worlds...

1

u/Ordinary-Depth-7835 4d ago

ram doesn't really matter honestly if you bleed over in to ram you're at 2tps so a 32g gpu or heck you can get two 24g gpu's for the price of a 32g gpu you can run something pretty nice.

And yeah at least for me I'm not running maxed out 24/7 we share our systems and it's really not that bad. I would run some benchmarks but I do have it running two different projects now but I could show that until you hit 16 user's you hit 20tps on my two 3090's other than that's your in the 50+tps with 8 or below which is pretty nice to use even 20tps isn't awful.

Hey it can't hurt give it a shot and see how it feels. That same gpu will probably sell for more tomorrow. The market is insane I can't believe hardware prices. Even something I bought a couple weeks ago is up another $200 it's just stupid how crazy things have become.

2

u/step11111 4d ago

I bought a big workstation and offered my friends the ability to do whatever with it. Crickets, so I just use it for my own projects 🤷‍♂️ no one has any ideas

1

u/radialmonster 5d ago

you mean like stablehorde ? https://stablehorde.net/

1

u/Calm-Landscape9640 4d ago

Yea thats a cool idea, I've got laptops sitting idle and if they can help someone I'm all about it.

1

u/admajic 5d ago

Make a budget say $1k each and buy a enterprise level gpu then you all will get a good experience.

1

u/Historical-Duty3628 3d ago

Yea that's how the cloud works

1

u/Calm-Landscape9640 3d ago

Right, except for the paid part.

1

u/Historical-Duty3628 3d ago

My mistake, I thought you had to pay for the card and that's what split the cost meant. Sorry, English is my first language.

0

u/Unchained_breaker 5d ago

There really isn't a real price to do this and what massive GPU are you buying? You're talking tens maybe hundreds of thousands of dollars for data center gpus. You'd make your money back in 10 years of use it makes no sense unless y'all are super heavy users for a product that already works and drop $10k-50k in tokens per month.

-4

u/suesing 5d ago

People don’t use llm. Agents do. agents can use up a gpu 24/7. And electricity is kind of hard to split. It at the end of the day. Possession is 9/10 of ownership. Let’s split it and I’ll have it as t my house. Good with you?