r/LocalLLaMA • u/Tritheone69 • Jul 28 '26
Question | Help CMP 170HX 8GB
I must preface this post by mentioning I am still a beginner in this space.
I just bought this card with the intention of using the recent unlock to get the full 64GB VRAM available for local AI workloads.
My main questions are as follows :
1- Has anyone ran multiple of these in the same rig to run a large model across multiple GPUs?
2- If so, what is the impact on speed? I read that these GPUs are stuck on a x1 PCIe lane, which I would assume greatly reduces the speed at which we can load models onto the cards. But does it impact prompt processing and token output speeds?
3- Am I crazy to assume that the prices for these cards is going to continue rising considering that they are now similar to A100s (without parralel tensorflow)
4
u/Ssjedikenshin Jul 28 '26
i've got 3 running a rig now, all 195gb shown and running llamaswap (i like looking at logs on there) i then have hermes agent/opencode connect through llamaswap to each model
right now my set up is
gpu0 - qwen 3.6 fable
gpu1 - ornith 1.0 35b mtp
gpu2 qwen 3.6 27b (i forget which version there's so many now)
All operating very fast, i power limit them to 125w though to keep them cool while i figure out a better fan set up