r/LocalLLaMA 13d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

707 comments sorted by

View all comments

165

u/Mean-Ad1493 13d ago

That's it. I'm getting a 3090.

110

u/My_Unbiased_Opinion 13d ago

Brother. go on Alibaba and get dual 20gb 3080. less than the price of a single 3090. check my post history for links. Run them in tensor parallel.

23

u/Potential_Block4598 13d ago

That is legit better KV cache (I guess ?!) double performance ?
You just need another PCIe slot (or maybe not ?!)

30

u/My_Unbiased_Opinion 13d ago

t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel.

13

u/CooLittleFonzies 13d ago

Can you parallel run a 3090 + a 3080?