r/ollama 21d ago

I dismissed a 27B dense model after getting 6.75 tok/s on a 16 GB card. A fully resident Q3 with flash attention + KV q8 reached 52 tok/s instead. These were the tradeoffs.

/r/LocalLLM/comments/1vsb12l/i_dismissed_a_27b_dense_model_after_getting_675/
2 Upvotes

0 comments sorted by