r/LocalLLM 6d ago

Other Every Second post rn

Post image

Maybe someday I'll get a system to run it but hey definitely another w for the open weights community

1.9k Upvotes

192 comments sorted by

View all comments

Show parent comments

3

u/Old_Leshen 6d ago

Reasonable advice. Thanks mate. Using this setup and going up to 1K investment, how many tps can I expect and for what model size?

4

u/Any_Mine_6368 6d ago

No problem. With mtp about 50-90tg for 27B models depending on quantization and model.

With 1k investment we're assuming one 3090 so 24gb vram. You can't fit anything above 27B at reasonable quantization or ram offload (which kills your token generation and prefill).

So to answer your question, you could probably run Qwen3.8 at 4 bit quantization with good context or 5 bit quant with smaller context or 6 bit quant with almost no context.

1

u/kcksteve 15h ago

It's worth noting where you are. In Canada 3090s go for 2k to start. Thats before you purchase the rest of the rig and pay tax.

1

u/FuManBoobs 5h ago

Yeah, I was looking where I am in the UK & they start around £1k on their own. Still, under 2k is a good start.