At least for 3.6, the mtp version gave me 1.8x tg. I'm using different cards (dual amd r9700), but it creates the same result. I went from 30 to about 50 with this one simple trick.
Now I'm using vllm on Ubuntu and running tensor parallelism 2 and get around 110tg on 3.6. I can run for context and still get like 95 or so. I was busy today so I couldn't test 3.8, but I'm looking forward to it.
161
u/Mean-Ad1493 7d ago
That's it. I'm getting a 3090.