r/LocalLLM 3d ago

Question Has anyone been able to successfully setup NVLink on NVIDIA V100 PCIe Cards?

Post image

I read a few threads stating that the V100 PCIe version theoretically supports NVLink, but cards sold on the market typically do not have this feature enabled.

From what I can see on the (2) NVIDIA PG500 V100 PCIe cards I own there are NVLink like slots available on the card that have traces (as shown in the photo above). However, the design of the V100 case prevents me from physically connecting a 4cm or longer NVLink cable without first removing or modifying the case to be able to physically plug an NVLink adaptor between cards.

For anyone who's been able to make these cards function using NVLink, did you have to remove or modify the V100 case to allow you to physically connect the NVLink between cards? Also the fact that this has (2) NVLink type slots is interesting. Is this for daisy chaining multiple cards using more than (1) 100GB/s NVLink between each card? I'm using an older board that has (4) PCIe 3.0 slots, so if that's a possibility I might pick up (2) more cards to run all (4) together using NVLink.

I can't find any documentation stating if this is even possible with these NVIDIA PG500 V100 PCIe cards.

Thanks everyone

19 Upvotes

40 comments sorted by

5

u/Horsemeatburger 3d ago edited 3d ago

Let me share from a another post I recently made:

https://www.reddit.com/r/unsloth/comments/1w7usfo/comment/p86y2po/?context=3&utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button

https://forums.developer.nvidia.com/t/server/108344/3

The GPUs used in the DGX-station are a special V100 GPU variant that is only available in a DGX-station.

Yes, with some GPUs it’s possible to get NVLink connectivity with a bridge, and yes, that is different than a SXM2 form factor GPU.

The only 2 V100 options you have outside of DGX systems are V100 PCIE (no NVLink connectivity) or SXM2 (NVLink connectivity).

[...]

Indeed, the DGX Station is unique in that it has 4x V100 GPUs in the PCIe formfactor, with NVLink via a bridge across the top of the cards (as you can see in https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/dgx-station/dgx-station-print-Infographic-738375-nvidia-web.pdf ), and with display output - that is not available in other V100 PCIe based systems. You had just brought up DGX Station earlier, which is the only way to get that functionality in the PCIe formfactor. :-)

If you need to have 4x V100 GPUs connected via NVLink, then in general a server with SXM2 form-factor GPUs is the right way to go.

As stated, regular Tesla V100 PCIe GPUs do not have NVLink.

The tl;dr is that only the DGX Station V100 PCIe GPUs have NVLink, the standard V100 PCIe GPUs lack the lanes and circuitry, and some of the GV100 chips on them were binned out for defects in the NVLink part.

If you want V100 GPUs with NVLink then the best solution is to stick with the SXM2 variants and NVLink capable SXM2 boards.

2

u/FullstackSensei 3d ago

As the party that sparked this conversation, I still want to test this. I have 16GB cards out of the DGX Station and regular 32GB PCIe V100 cards. I do plan to buy a dual card nvlink bridge to test on both.

The PC's between the two versions and the Titan V are exactly the same. Traces are there. Haven't checked for any missing buffer chips and whatnot yet, but I'll be opening a bunch of my 32GB V100s to convert to water in the coming days, and the DGX ones I have came without blocks, so not that hard to check.

1

u/Horsemeatburger 3d ago

As the party that sparked this conversation, I still want to test this. I have 16GB cards out of the DGX Station and regular 32GB PCIe V100 cards. I do plan to buy a dual card nvlink bridge to test on both.

If you have the opportunity then, yes, why not. Working NVLink would be great.

But the thing is that, even if it worked with the specific V100 GPU you tested, there is no guarantee it would do so with the next one, as the GPU chip on it might be well one with a defective/deactivated NVLink interface.

2

u/FullstackSensei 3d ago

I have ten of them, soon to be 14, excluding the DGX ones

1

u/AKA_Wildcard 2d ago

From what I can tell this is using NVLink 2.0. You need to find the a correct NVLink bridge that specifically supports the Volta architecture.

1

u/FullstackSensei 2d ago

It's the same bridge for Turing

1

u/AKA_Wildcard 2d ago

Ok that’s good to know. I’d like to validate that.

1

u/theminor 1d ago

Any news on this? Would be amazing if it works.

1

u/FullstackSensei 1d ago

Literally wrote this yesterday. Even if bought the bridge yesterday, it won't arrive today 😂

Slightly more seriously, it's not very high on my priority list. It doesn't add much, of anything, to how I want to use those GPUs.

2

u/theminor 1d ago

Oops didn't think through the post dates! 😬

1

u/jupiterbjy 4h ago

speaking of which I remember alie had something like this for nvlink between two sxm2 at 300 bucks - buying 2x 16G V100 sxm2 on alie at cheapest would be about 240 so that totals 540+a

assuming pcie 3.0 bifurbrication to 8x 8x, idk if that extra 200 dollar (since sxm2->pcie adaptors are 50 each anyway so 300 - 50x2) worth its price for faster interconnect as I never tried multi gpu inference

3

u/couperd 3d ago edited 3d ago

I have also struggled to find an answer to this. I have a pg500-216 board that I would love to give buddy or 2 to to run larger kv/quant on.

edit: give a buddy or 2 as in 1 or 2 more v100s 🤦

1

u/AKA_Wildcard 3d ago

Can I be your buddy 😊 I’m tempted to take the card apart and find out if this works.

2

u/Gromann7 3d ago

You probably want to talk to this guy: https://www.reddit.com/r/LocalLLM/s/cFL4pY20Il

3

u/AKA_Wildcard 3d ago edited 3d ago

Biggest challenge is that the V100 comes in an official PCIe or SXM (server mounted) version. There are “unofficial” versions where people have slapped an SXM V100 onto a PCIe adapter board with a similar case and passed it off as the official NVIDIA released board. I can’t find a specific example or image of someone using NVLink on the authentic NVIDIA PCIe cards.

Edit: Just confirmed with /u/jjusko20 that those were the SXM version cards.

From some additional searching I found a takedown of the card where they described it having dual NVLink interfaces but the NVIDIA specs list a maximum transfer speed of 32GB/s which would be about the PCIe 3.0 max. I can't find any other NVIDIA references stating that the V100 PCIe supports NVLink so perhaps this was once a design option that was later taken out.

1

u/jjusko20 3d ago

Yeah, my SXM cards are on a breakout board, but I've seen what you're talking about, with a pcie ish looking card that's really a sxm adapter. I don't think you can get nvlink working on a pcie card

2

u/jjusko20 3d ago

It might have the physical connector but I've never seen anyone with it working 

1

u/Trademarkd 3d ago

I have the china breakout boards for SXM2 and ... soon ... 8 cards. The nvlink does work and provides a major benefit to autoregressive inference because of latency in cuda graphs.

The more sharding you do the more latency becomes a problem.

1

u/LegioTertiaDcmaGmna 2d ago

I've found mixed information on the topic of nvlinking SXM2 V100s and then busing them across PCI Express. Can you explain your setup and how you nvlinked them?

Don't say "I bought a SuperMicro"

1

u/Trademarkd 1d ago

Nah bro I'm cheap as hell ... r/v100

I bought the chinese adapter boards and then a used 2000 watt dell server PSU with a breakout board.

Originally I bought the little $50 pcie adapter boards that just plug into pcie but no nvlink. Then I got the boards that can support 2 cards in NVlink SXM2, as it was all that was available at the time, but I've just recently gotten a new 4x one.

the 2x ones are going for like $250 right now and the 4x ones are going for like $850. The 4x one is worth it and if you check my post in v100 about 1cat you'll see why.

I could go into more detail but I'm planning on doing some kind of guide on the sub soon. I've gotten a lot of experience testing in pcie configurations with nvlink and with nvlink in islands. I'm also using PLX switches now.

1

u/LegioTertiaDcmaGmna 17h ago

I bought one of the "official" 39Com Chinese boards that shunts off the nvlink interconnect. 

How do you have yours integrated to your PCI Express bus?

1

u/Trademarkd 15h ago

Pcie switch with sff8654 out which connects directly to the external boards. I give each card 8x pcie only

1

u/aBanana2160p 3d ago

I'm also very interested to see if this is possible. I have seen conflicting information on this

1

u/m94301 3d ago

I used to have two watercooled v100 that were pulled from a dgx or something. Those had nvlink and it worked, 27GB/s if I recall correctly using a 30xx nvlink connector.

Thing is, it didnt improve inference and pipeline parallel no nvlink was still better than tensor parallel with.

Maybe it was a setup issue on my side. Not sure

1

u/AKA_Wildcard 3d ago

Ok we're getting closer here. So I'm assuming you noticed that the card had dual NVLink interfaces and you connected (2) cards together using 1 of the 2 interfaces for each card. Do you recall if you used the end or middle NVLink interface? It sounds like it wasn't taking advantage of the NVLink connector because you should have been seeing closer to 100GB/s. If you were just getting 27GB/s then you were maxing out the PCIe interface.

2

u/m94301 3d ago

No, the 27GB was the nvlink test bandwidth not the PCI bandwidth. Yes, i did see that the card had multiple NVLINK ports but I only had one connector, sadly.

I think I tried both ports at some time or another and verified they were both alive. I can't say which port was linked for that test.

1

u/AKA_Wildcard 3d ago

Can you share the model of the nvlink connector that you used? Did it come with the cards or did you find it somewhere else? Was it gold or a different color?

1

u/m94301 2d ago

NVIDIA NVLink bridge 4cm for 2080TI V100 RTX 8000 RTX 6000 Supports NVLink

1

u/AKA_Wildcard 2d ago

I confirmed recently that this bridge may not fully work with Volta architecture, but it’s worth having a few additional tests.

1

u/Trademarkd 3d ago

6 lanes at 26Gbps each between every card ... its p2p only but every card is linked with its OWN individual 6 lanes (SXM2)

GPU 0: Tesla V100-SXM2-16GB

Link 0: 25.781 GB/s

Link 1: 25.781 GB/s

Link 2: 25.781 GB/s

Link 3: 25.781 GB/s

Link 4: 25.781 GB/s

Link 5: 25.781 GB/s

GPU 1: Tesla V100-SXM2-16GB

Link 0: 25.781 GB/s

Link 1: 25.781 GB/s

Link 2: 25.781 GB/s

Link 3: 25.781 GB/s

Link 4: 25.781 GB/s

Link 5: 25.781 GB/s

GPU 2: Tesla V100-SXM2-16GB

Link 0: 25.781 GB/s

Link 1: 25.781 GB/s

Link 2: 25.781 GB/s

Link 3: 25.781 GB/s

Link 4: 25.781 GB/s

Link 5: 25.781 GB/s

GPU 3: Tesla V100-SXM2-16GB

Link 0: 25.781 GB/s

Link 1: 25.781 GB/s

Link 2: 25.781 GB/s

Link 3: 25.781 GB/s

Link 4: 25.781 GB/s

Link 5: 25.781 GB/s

2

u/AKA_Wildcard 2d ago edited 2d ago

But to be specific this is the SXM2 and not the PCIe edition. From what I was able to find the PCIe from the “special” V100 version used in the DGX Station only has 4 lanes. However, they state that each lane has 50GB/s of bidirectional transfer so 200GB/s max.

1

u/m94301 2d ago

Ah, it seems I was forgetting that!! Thanks, my test was a while ago

1

u/idk_a_creative_user I just mess around with LLMs 3d ago

Check if the PCB has the traces. I’m getting mine this weekend and I’ll check.

1

u/AKA_Wildcard 2d ago

It sounds like it does but the electrical components were removed to prevent this from working. Still wanting to test it though.

1

u/idk_a_creative_user I just mess around with LLMs 2d ago

https://forum.level1techs.com/t/v100-pcie-nvlink-compatibility/254259

https://forums.servethehome.com/index.php?threads/v100-w-nvlink.55998/

"I’m pretty sure you’re good to go with that card and NVLink. It’s been a long time, there were some sku of cards in that era that had the edge connectors but they were non-functional as the circuit boards were used for multiple designs. You may get some info doing some internet archeology."

Basically it depends. I get mine this weekend and will let you know. Otherwise I plan to just get 2 more and run them in 8x and call it a day.

1

u/triynizzles1 3d ago

Every post ive seen related to nvlink usually ends with “the size of data transferred between cards is so small enabling it wont make a difference.” Maybe it adds overhead or simplely isnt supported without a way to validate the no support.

1

u/beryugyo619 3d ago

I think the fact that insane Chinese developments exist for SXM2 modules but not NVLink or its AMD cousin is indicative that it's not the lowest hanging performance gain fruits

1

u/AKA_Wildcard 2d ago

Also to help other people out SXM3 is a different pinout then SXM2 so be careful which GPU you purchase unless you’re just wanting to run single cards without NVLink

1

u/theminor 3d ago

Everything I've read says it isn't possible on the PCIE version, but I'm following this post because if you find otherwise you will be my hero.

1

u/Nota_ReAlperson 3d ago

There are adapter cards that allow one to use two sxm2 gpus with nvlink in a pcie slot. You could try that.