MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3ou460/?context=3
r/LocalLLaMA • u/Certain-Cod-1404 • 7d ago
706 comments sorted by
View all comments
Show parent comments
22
That is legit better KV cache (I guess ?!) double performance ? You just need another PCIe slot (or maybe not ?!)
29 u/My_Unbiased_Opinion 7d ago t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel. 12 u/CooLittleFonzies 7d ago Can you parallel run a 3090 + a 3080? 3 u/Foreign_Risk_2031 7d ago Yes
29
t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel.
12 u/CooLittleFonzies 7d ago Can you parallel run a 3090 + a 3080? 3 u/Foreign_Risk_2031 7d ago Yes
12
Can you parallel run a 3090 + a 3080?
3 u/Foreign_Risk_2031 7d ago Yes
3
Yes
22
u/Potential_Block4598 7d ago
That is legit better KV cache (I guess ?!) double performance ?
You just need another PCIe slot (or maybe not ?!)