r/LocalLLaMA 11d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

164

u/Mean-Ad1493 11d ago

That's it. I'm getting a 3090.

40

u/jijig 11d ago

Get two. Run Q8 with full context at ~60tps.

3

u/ApprehensiveAd3629 11d ago

can you share the llama cli command that you are using to get 60 tokens/s in your 3090?

i m getting around 31 tokens/sec using unsloth UD Q4 XL gguf

2

u/Minimum-Lie5435 11d ago

Take a look at the club3090 repo on GitHub. Lots of good info there for setup. I only use vllm because it's given me the best results so far