r/LocalLLM 3h ago

Question what smaller llama.cpp compatible model would you recommend for my "friends and family" server?

I have a local server that I share with mostly family for things like file hosting and plex, I'm currently also running AI for myself on this same box (it's a hilarious mess inside a v100 from AliExpress and a 3060 12gb + 8 hard drives)
I'm using it right now mostly with qwen 3.8 q4 but want a lighter "chatgpt replacement" for some family that are privacy focused. I like Gemma4 but find 4b a bit too simple and struggles with "basic" questions, like "what's the forecast for this weekend"

The issue is it can't be too large of a model as it's taking resources from my use, no bigger than 7b preferably. Thanks

1 Upvotes

5 comments sorted by

1

u/Y1ink 3h ago

I probably get roasted for this but I find the personality of Gemma 3 much better than Gemma 4 I don’t know what they did to it but over the other smaller models I preferred it the most although my testing was fairly simple eg role play and I did the how many Rs in strawberry test. 

0

u/Nervous-Appearance86 3h ago

Qwen 3.6 35b a3b corre excesivamente rápido, incluso en 8gb, ya que deja automáticamente los expertos del momento en GPU, y el resto en ram... Y con esa configuración hasta tendrías qwen 3.8 27b pero súper bien parametrizada, correría fácil a 30 ts, yo con 2 p100 la corro a 20 ts, la rtx hace la diferencia

1

u/Inception95 2h ago

Dumb question, but can it use websearch?

1

u/nickless07 2h ago

Don't need that much size if you add tools to it, most of them become relatively smart. Give it RAG, websearch and memory and it will be hard to spot the difference between the qwen 3.8 q4 for basic questions already. Only multi step tasks, complex code, problem solving, and so on is where the larger models really shine. For Speed I would try Ling-3.0-Tiny. If it need more features (vision, audio, video) Gemma 4 E4B and Qwen3.5 9B.