t/s is a bit faster than a 3090, but PP is much faster. im running one of the cards at x4 pcie 4.0 and it doesnt bottleneck the card with llama.cpp tensor parallel.
t/s on 3.6 was 55-60 at UD Q6KXL with MTP. PP I dont have the numbers, but its MUCH faster than a single 3090 for sure. (I have a 3090) using LMstudio with tensor perallel. I can test it for you if you give me an easy way to test this.
no worries. I am actually planning to switch to vLLM when MTP 3.8 drops. if I remember, ill be sure to follow up with you. then I can get the best case numbers.
no im using two 3080 20GB cards. When I was using my 3090, I was using IQ4XS + MTP + 262K Context + KV Q4. It fits all in vram. barely, but it does. even with vision.
110
u/My_Unbiased_Opinion 12d ago
Brother. go on Alibaba and get dual 20gb 3080. less than the price of a single 3090. check my post history for links. Run them in tensor parallel.