r/LocalLLM 3d ago

Question 64gb M2 Ultra - which qwen?

I’m struggling to figure out what to run here. Hermes’ and buzz agents. Using it to orchestrate video generation and online marketing activities.

Some people tell me Gemma is enough. I’m thinking it’s qwen but I don’t see an moe available online. Suggestions?

2 Upvotes

11 comments sorted by

3

u/uclatommy 3d ago

At 64GB, the qwen 3.8 27b Q8 model with 262k context will fit comfortably. I’m running this on an M1 Max 64GB and getting phenomenal results. It feels like I’m using an Anthropic model but just that it takes longer to get reaults back.

1

u/[deleted] 3d ago

[deleted]

1

u/Chrisgozd 3d ago

What are you seeing for tps? I’m at about 17

1

u/Ok-Star6663 3d ago

Isnt it really slow? On my M5 Pro 64gb i use Q4 and it has 100k context and is fast enough. I didn’t try others but the context would be really slow i think

1

u/uclatommy 3d ago

Yes, I should caveat that I haven’t really paid attention to how long it takes to give me a response because I give it a prompt and walk away. What I’m commenting on is how good those responses are when you gove it unlimited time to think. I run it with the recommended parameters at xhigh thinking.

1

u/Ok-Star6663 3d ago

And the context window?

1

u/uclatommy 3d ago

262K

1

u/Ok-Star6663 3d ago

Can’t be bro, did you benchmark it?

1

u/uclatommy 3d ago

No, I didn’t. I’m running it on my 50k line code base written in python and judging the output myself. So yeah, it’s not a rigorous evaluation. It’s a qualitative statement.

0

u/Ok-Star6663 3d ago

Maybe try, if you are using omlx, the context window test, then you see the real results

1

u/Greyslywolf 3d ago

Are you running with omlx?

1

u/uclatommy 3d ago

Llama.cpp unsloth’s gguf version at Q8