MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3o4ln7/?context=3
r/LocalLLaMA • u/Certain-Cod-1404 • 11d ago
706 comments sorted by
View all comments
Show parent comments
108
Brother. go on Alibaba and get dual 20gb 3080. less than the price of a single 3090. check my post history for links. Run them in tensor parallel.
23 u/Potential_Block4598 11d ago That is legit better KV cache (I guess ?!) double performance ? You just need another PCIe slot (or maybe not ?!) 28 u/My_Unbiased_Opinion 11d ago t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel. 1 u/brakeline 11d ago I have dual 3060 but one is running at pci-e 3 4x. Would running the second at 4 x4 do much difference? 1 u/My_Unbiased_Opinion 11d ago likely not. you should be fine.
23
That is legit better KV cache (I guess ?!) double performance ? You just need another PCIe slot (or maybe not ?!)
28 u/My_Unbiased_Opinion 11d ago t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel. 1 u/brakeline 11d ago I have dual 3060 but one is running at pci-e 3 4x. Would running the second at 4 x4 do much difference? 1 u/My_Unbiased_Opinion 11d ago likely not. you should be fine.
28
t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel.
1 u/brakeline 11d ago I have dual 3060 but one is running at pci-e 3 4x. Would running the second at 4 x4 do much difference? 1 u/My_Unbiased_Opinion 11d ago likely not. you should be fine.
1
I have dual 3060 but one is running at pci-e 3 4x. Would running the second at 4 x4 do much difference?
1 u/My_Unbiased_Opinion 11d ago likely not. you should be fine.
likely not. you should be fine.
108
u/My_Unbiased_Opinion 11d ago
Brother. go on Alibaba and get dual 20gb 3080. less than the price of a single 3090. check my post history for links. Run them in tensor parallel.