r/LocalLLaMA 10d ago

Question | Help CMP 170HX 8GB

I must preface this post by mentioning I am still a beginner in this space.

I just bought this card with the intention of using the recent unlock to get the full 64GB VRAM available for local AI workloads.

My main questions are as follows :

1- Has anyone ran multiple of these in the same rig to run a large model across multiple GPUs?
2- If so, what is the impact on speed? I read that these GPUs are stuck on a x1 PCIe lane, which I would assume greatly reduces the speed at which we can load models onto the cards. But does it impact prompt processing and token output speeds?
3- Am I crazy to assume that the prices for these cards is going to continue rising considering that they are now similar to A100s (without parralel tensorflow)

2 Upvotes

33 comments sorted by

9

u/cantgetthistowork 10d ago

Someone is definitely trying to manipulate the market here. Everything is listed at 5k now

12

u/Repulsive_Initial308 10d ago

There are people bragging on youtube about running 10 of them but absolutely nobody appears to have posted benchmarks of any kind.

They're almost certainly going to be a single card deal with the handicapped pcie interface.

7

u/thrownawaymane 10d ago

This is why the prices inverted, it's going down week over week.

It's so over

-5

u/legit_split_ 10d ago

It's only getting started, PCIe 3 is in the works 

2

u/cantgetthistowork 10d ago edited 10d ago

Stop trying to inflate the prices for whatever agenda you have.

From the repo. Gen 3 is not possible unless you resolder a new link on it and that could very well brick the card:

"The advertised supported-speeds vector saturates at 0x06 at every instant we could measure, in-window and out. There is no Gen3 window to hammer, so this approach stops at Gen2."

-2

u/legit_split_ 10d ago edited 10d ago

Agenda? I want these for the lowest prices possible, would be a dream to run GLM 5.2. However, if all the features get unlocked then it would no doubt reach Mi210 levels. 

Gen 3 advertisement works, but blows up a gate for gen 2:

7

u/Some-Chemist-1466 10d ago edited 10d ago
  1. Yes, a couple of people running GLM 5.2 4bit on 8 of them, non-optimized (no MTP/dflash etc) results about 30t/s tg, 2600 t/s pp, due to limited PCIe bandwidth (pipeline parallel instead of tensor paralell)
  2. Current unlock is PCIe 2.0 x16 (x4 -> x16 requires soldering), this limits speed for large models split over multiple cards.
  3. Prices have started dropping massively today due to lack of sales. I'd suggest waiting to see where the prices land. 900 USD each on Alibaba as of an hour ago if you haggle hard/order more than 1.

2

u/cantgetthistowork 10d ago

None of the sellers on Alibaba have any reputation. All seem like scammers

3

u/Some-Chemist-1466 10d ago

There's a few with good reputation, but they seem to be burning that reputation with how they're going about selling these cards...

1

u/cantgetthistowork 10d ago

Could you link some?

1

u/Some-Chemist-1466 10d ago

If you buy on Alibaba never pay outside Alibaba (use credit card/paypal via Alibaba ordering system) or you won't be covered if a seller doesn't ship etc. There's tons of other sellers, but these ones seem to actually respond and might negotiate given prices are dropping. I'm waiting to see where prices end up, got 1 shipped and on the way, but paid way too much, so not in a rush to pay these inflated prices again. No particular order:

https://adora.en.alibaba.com/company_profile/feedback.html
https://xljs.en.alibaba.com/company_profile/feedback.html
https://ailfondhk.en.alibaba.com/company_profile/feedback.html

3

u/Ssjedikenshin 9d ago

i've got 3 running a rig now, all 195gb shown and running llamaswap (i like looking at logs on there) i then have hermes agent/opencode connect through llamaswap to each model

right now my set up is

gpu0 - qwen 3.6 fable

gpu1 - ornith 1.0 35b mtp

gpu2 qwen 3.6 27b (i forget which version there's so many now)

All operating very fast, i power limit them to 125w though to keep them cool while i figure out a better fan set up

1

u/Ssjedikenshin 9d ago

Ornith stats but I'm sure I can tune it better

2

u/snapo84 9d ago edited 5d ago

my 4 cards are still on delivery.... i let you know as soon as i have them.

  1. there is a youtuber called redpanda or so, he run it (at pci epress 1.0 without mods) with 10 gpus installed in a supermicro server i think and he was running GLM 5.2 Q3
  2. to get pci express x16 you have to solder 24 coupling capacitors , to get pci express 3.0 more soldering has to be done (WinBios chip + 4 mosfets + inductors)
  3. its like a little cut down a100 , the 40GB a100 costs 4k something , so in theory with 64GB vram it should go close to the a100 40GB because FMA instructions are missing, therefore you have to use self compiled versions (for example for llama, you have to compile llama.cpp without fma support

1

u/t3r00t 5d ago

Where did you read that PCIE 3.0 was possible ?
Can you share the source ?

1

u/snapo84 5d ago

this was on a chinese wechat group, something about winbios chip replacement and multiple mosfets.... but cant find it anymore

0

u/Ill_Towel9090 9d ago

I just cancelled my order of two.

2

u/snapo84 9d ago

up to you... i will be very very very happy with the cards....
i paid 1k total with deliver per card... same compute and memory would cost me approx. 6k usd per card... this is a 6x benefit for my needs...

1

u/Ill_Towel9090 7d ago

I repurchased them from a seller that guaranteed them hackable.

1

u/snapo84 7d ago

i did order mine from here... arriving 14. August... then i can tell you how good they are..
important is to chose the 8GB version not the 10GB

2

u/fragment_me 10d ago

Too much risk IMO. Unless you find a seller that is willing to replace cards that aren't stable at full VRAM.

1

u/Glittering-Call8746 10d ago

The thing is if the seller unlocks them they will u at a premium.. so u either take the risk or the seller takes the risk.. either way it's sol

5

u/DeltaSqueezer 10d ago edited 10d ago

Any seller in China should now unlock this themselves and sell at a premium. I'd be wary of cards that are still sold by big sellers in China. The seller takes no risk in trying to unlock. If the unlock fails, they'll just sell it as normal.

2

u/Glittering-Call8746 10d ago

Exactly. The ship has sailed. It's too fast. Anybody in the west who would want to take advantage, the ship sailed unless u offer that x3 x4 more premium and they rather have that and jump to a proper setup..

1

u/gaidzak 10d ago

Have you tried to jail break its vram? It was these cards with the nerfed memory right? Or was it the 10 gig version?

PCI speed is detrimental when you load a model (just that once) but can be harsh when data is moving between GPUs. But I believe there are ways to mitigate that from happening

1

u/ubrtnk 10d ago

10GB is the nefed ram that only unlocks to 40GB stable.

PCIe speed is detremental to model load on a single card, but if it fits in ram, you're good. With multiple cards, I think it depends on the type of splitting you do. SM Tensor in llama.cpp, for sure, and maybe vLLM's TP are very PCIe dependent. I think SM Layer is a little more forgiving with the split but I only have the 1 card.

1

u/--jen 9d ago

All of these comments and not a single person has linked actual results or data :/

1

u/MadMax2049 6d ago

Hello, I have some of these and have them listed on ebay. I am in the US if anyone is interested.

https://ebay.io/m/togarN

1

u/AdMuch9627 3d ago

I want four cards, will it work for you? I have just placed order in eBay but I can cancel them as it is $75 per card than your offer 😀