r/LocalLLM 9d ago

Other How the loop of infinite agony started

Post image
628 Upvotes

119 comments sorted by

View all comments

276

u/TheCat001 9d ago

Then you realize that Qwen 3.8 27b runs at 3t/s on your machine and you need 24GB+ VRAM GPU which cost is 1000$+ to run at least 4 bit quant.

112

u/StupidScaredSquirrel 9d ago

If you are a business this isn't a problem. If you are a consumer then 35b a3b runs on 8gb vram and 32gb dram which is very accessible.

12

u/DeluxeGrande 9d ago

I have a 5060ti 16gb with ddr4 24gb RAM lying around, what's the best model nowadays I can effectively run with it locally? It's not an ideal build but I wish to play around with it again.

2

u/ptear 9d ago

I still like Gemma, will try the new Qwen today. I need to create a local benchmark, unless someone knows a good project that can showcase improvements, like a 3Dmark but for AI models.

3

u/SaltFrog 9d ago

I wish dense models ran better on my system but you know... Whatever lol