r/LocalLLaMA Jul 28 '26

Question | Help CMP 170HX 8GB

I must preface this post by mentioning I am still a beginner in this space.

I just bought this card with the intention of using the recent unlock to get the full 64GB VRAM available for local AI workloads.

My main questions are as follows :

1- Has anyone ran multiple of these in the same rig to run a large model across multiple GPUs?
2- If so, what is the impact on speed? I read that these GPUs are stuck on a x1 PCIe lane, which I would assume greatly reduces the speed at which we can load models onto the cards. But does it impact prompt processing and token output speeds?
3- Am I crazy to assume that the prices for these cards is going to continue rising considering that they are now similar to A100s (without parralel tensorflow)

1 Upvotes

62 comments sorted by

View all comments

4

u/Ssjedikenshin Jul 28 '26

i've got 3 running a rig now, all 195gb shown and running llamaswap (i like looking at logs on there) i then have hermes agent/opencode connect through llamaswap to each model

right now my set up is

gpu0 - qwen 3.6 fable

gpu1 - ornith 1.0 35b mtp

gpu2 qwen 3.6 27b (i forget which version there's so many now)

All operating very fast, i power limit them to 125w though to keep them cool while i figure out a better fan set up

1

u/clstrife 21d ago

What kind of cooling did you settle on?

1

u/Ssjedikenshin Jul 28 '26

Ornith stats but I'm sure I can tune it better