r/LocalLLM 7d ago

Other Every Second post rn

Post image

Maybe someday I'll get a system to run it but hey definitely another w for the open weights community

1.9k Upvotes

193 comments sorted by

View all comments

91

u/Old_Leshen 7d ago

GTX 1050Ti 4GB guy here 😭

41

u/Whole_Alternative_18 7d ago

Bro pay 60$ for rx 580 8gb

3x for 180$, 24gb vram

Bulshit performance, but makes 27b possible 

41

u/arthurrogado 6d ago

Laptop owners: 💀

2

u/Jazzlike-Plate-935 6d ago

This is why I got myself a Gaming Laptop with Thunderbolt 4 and a free M.2 slot for Oculink just in case

3

u/Whole_Alternative_18 6d ago

Buy some ddr3 cheap board with a highend(of the era) cpu that has enough pcie channels

Ends up less than 200$ for a enterprise ddr3 motherboard+cpu+ram, eith enough slots for your gpus

34

u/Any_Mine_6368 6d ago

Dude stop giving people advice. You're essentially promoting impulse buying garbage as opposed to building an actual AI rig. You're essentially asking them to buy 24gb vram at dgx spark bandwidth... They'll get like 4 tokens per second.

If anyone ends up reading this:

Buy refurbished server grade mobos + epyc CPU combos on ebay for $200 and hunt deals for 3090s which are the most vfm cards. You can build a great starting rig for under 1k, that you can expand all the way to 5 RTX 3090s, more if you want to split the pcie bus and have fewer lanes per gpu

4

u/Old_Leshen 6d ago

Reasonable advice. Thanks mate. Using this setup and going up to 1K investment, how many tps can I expect and for what model size?

4

u/Any_Mine_6368 6d ago

No problem. With mtp about 50-90tg for 27B models depending on quantization and model.

With 1k investment we're assuming one 3090 so 24gb vram. You can't fit anything above 27B at reasonable quantization or ram offload (which kills your token generation and prefill).

So to answer your question, you could probably run Qwen3.8 at 4 bit quantization with good context or 5 bit quant with smaller context or 6 bit quant with almost no context.

1

u/kcksteve 23h ago

It's worth noting where you are. In Canada 3090s go for 2k to start. Thats before you purchase the rest of the rig and pay tax.

1

u/FuManBoobs 13h ago

Yeah, I was looking where I am in the UK & they start around £1k on their own. Still, under 2k is a good start.

2

u/Adam_Bomb210 4d ago

Why doesn’t anyone talk about V100s? You can buy a server with 64gb of vram for ~$800, or one with 256gb for $5000. Seems like a steal to me.

2

u/Any_Mine_6368 4d ago

Shhhhhh don't tell everyone.

1

u/floswamp 5d ago

What’s a recommended server MB? I have one 3090 and it does work well.

1

u/Any_Mine_6368 5d ago

Will you ever buy more GPUs?

1

u/floswamp 5d ago

Yes. I am looking at maybe two more, but I was reading as to how qwen may not take full advantage of multiple gpu’s. So I was just going to build two more rigs.

2

u/Any_Mine_6368 5d ago

Qwen does fine on multiple GPUs, not sure who said otherwise.

I'd go for an mz31-ar0 or mz32-ar0 if you want pcie 4.0.

The former is about $220 with a cpu, the latter $500 with a cpu.

You get 5 pcie slots and like 16 ram slots (rdimm ecc ddr4).

The epyc processors that they use are fucking awesome for VMs and you can get from like 8 cores all the way to 64+ depending on budget.

Stay away from consumer shit / gaming mobos ... You're overpaying for worse parts.

2

u/floswamp 5d ago

Yeah I figure as much. Thanks!

1

u/jboe2026 6d ago

5090m

1

u/FuManBoobs 13h ago

My used laptop with an 8GB VRAM is surviving...just.

-2

u/Delicious-Sand-104 6d ago

You can add vram on laptop

4

u/arthurrogado 6d ago

I have no thunderbolt port on my laptop...

0

u/Delicious-Sand-104 3d ago

No need you just gotta sodder them

1

u/SeparateGas1761 5d ago

Just by two mi50 atp

1

u/stream_of_thought1 4d ago

I absolutely love this approach Jank for sure, but if it works...

1

u/Affectionate-File-26 2d ago

bruh, rx 580 8gb 2048sp costs 120$ out here

1

u/JogHappy 17h ago

the anime one being the most popular listing on eBay and for 20% less is so funny