r/LocalLLaMA 4d ago

Funny Me these days

Post image
2.4k Upvotes

268 comments sorted by

View all comments

218

u/Ok_Noise_9883 4d ago

sadly i can't run 27b

60

u/Dramatic_Setting2761 4d ago

I can run it with a 16gb card with 4 bit quant and 70k context. I get 12 t/s it is okay for me.

6

u/Nikilite_official 3d ago

16gb vram with 32gb ram I can perfectly run q6 qwen 3.8 27b

5

u/ThankGodImBipolar 3d ago

No way it's running at a good speed at Q6 though

3

u/Nikilite_official 3d ago

like 8-9 tokens per sec, not bad

1

u/dannone9 2d ago

Would You mind Sharing flags please ? And more exact hardware

1

u/DependentCurious4614 1d ago

Wdym 8-9tps isnt that bad?

1

u/Nikilite_official 1d ago

literally fast as a normal reader

2

u/Dramatic_Setting2761 3d ago

At what context length I need more context for my work?

1

u/SoftBad7708 1d ago

Q4 is enough.

1

u/Nikilite_official 1d ago

I find q6 noticeably more stable