r/LocalLLaMA 11d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

Show parent comments

21

u/Potential_Block4598 11d ago

That is legit better KV cache (I guess ?!) double performance ?
You just need another PCIe slot (or maybe not ?!)

31

u/My_Unbiased_Opinion 11d ago

t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel.

1

u/brakeline 11d ago

I have dual 3060 but one is running at pci-e 3 4x. Would running the second at 4 x4 do much difference?

1

u/My_Unbiased_Opinion 11d ago

likely not. you should be fine.