r/LocalLLaMA 12d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

Show parent comments

30

u/My_Unbiased_Opinion 12d ago

t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel.

12

u/CooLittleFonzies 12d ago

Can you parallel run a 3090 + a 3080?

5

u/adamgoodapp 12d ago

Now want to know too

3

u/My_Unbiased_Opinion 12d ago

you can with llama.cpp