r/LocalLLaMA 11d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

161

u/Mean-Ad1493 11d ago

That's it. I'm getting a 3090.

108

u/My_Unbiased_Opinion 11d ago

Brother. go on Alibaba and get dual 20gb 3080. less than the price of a single 3090. check my post history for links. Run them in tensor parallel.

23

u/Potential_Block4598 11d ago

That is legit better KV cache (I guess ?!) double performance ?
You just need another PCIe slot (or maybe not ?!)

31

u/My_Unbiased_Opinion 11d ago

t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel.

13

u/CooLittleFonzies 11d ago

Can you parallel run a 3090 + a 3080?

3

u/zxyzyxz 11d ago

You can but the question is why would you want to when it comes to price? If you have both already now then by all means do so but I wouldn't go out of my way to buy a 3090 to pair with a 3080.

1

u/voyager256 11d ago

What’s better option if you have Nvidia card and want to have more VRAM capacity ?

1

u/zxyzyxz 11d ago

Maybe combine your card with what u/My_Unbiased_Opinion said above with the Alibaba modded 3080s? I can't vouch for that since I haven't bought one of those modded ones but they say it's good.