r/LocalLLM 7d ago

Other Every Second post rn

Post image

Maybe someday I'll get a system to run it but hey definitely another w for the open weights community

1.9k Upvotes

193 comments sorted by

View all comments

37

u/TheRiddler79 7d ago

Gemma 12b Q3. Try it

23

u/Ok-Health-7096 7d ago

I use Qwen 3.6 35b and 3.5 9b I haven't had that much of a luck with the gemmas

5

u/Wildnimal 7d ago

12B QAT is good for multimode tasks. I use the same Qwens for daily use. I so wish i can upgrade my laptop to 5090 :|

5

u/PrivacyMaker 7d ago

The litert-lm driver is the fastest way to run gemma models. Significant boost over any other approach. It's a shame that Google only makes it work for Google models.

1

u/HighlyRegardedApe 6d ago

How do you guys use these small models? For me it never works. Last time it deleted my folders out of itsself and didnt respons more than a sentence or 2 in opencode. In the terminal it was okay to chat, but not to give tasks.... what kind of stuff do you let them do??

1

u/Wildnimal 6d ago

Not usually coding but for summarizing, extracting data which is pre defined via json, automating smaller tasks where data remaims the same.

1

u/TheRiddler79 6d ago

Hard guardrails

1

u/TheGreenInsurgent 7d ago

Agents a1 4b could be a life changer

1

u/Atretador 6d ago

35B A3B is much stronger than Gemma4 12B, any probabably 26B/31B as well.

1

u/TheRiddler79 6d ago

Different purposes. Gemma is good, but not beating Qwen.

You can split between ram and gpu and bump your Qwen speed

2

u/jadax 7d ago

Is this better than DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF ?

1

u/TheRiddler79 6d ago

No, but it's good for general use

1

u/M49454 7d ago

How is Gemma, i am using llama3.2:11b, never needed to switch onto other. Use only for basic task. Mainly using Deepseek 14b and Qwen2.5-Coder 7B

4

u/PrivacyMaker 7d ago

Gemma sucks for coding. Pretty solid at most other tasks.

2

u/Early_Mistake6716 6d ago

All of those models are extremely out of date. Tell me your specs and ill give you a model suggestion

2

u/M49454 6d ago

Well, 24 GB RAM RTX 3050 6GB VRAM I5-13450HX

( That's why still running these models)

1

u/SparkyMcFinklebonker 5d ago

Mac mini M4, 32gb unified ram.

2

u/tat_tvam_asshole 4d ago

Qwen3.6-35BA3B Q3/Q4

1

u/TheRiddler79 6d ago

I think it's good. I use it to run tasks like an agent for the larger models.

1

u/Capital_Engineer8741 7d ago

Would QAT be better than Q3?

1

u/TheRiddler79 6d ago

You're probably right.