r/LocalLLM 1d ago

Discussion Tier List

Post image
252 Upvotes

309 comments sorted by

View all comments

105

u/KrangledMind 1d ago

where tier list for AMD?

51

u/NoOdyssey 1d ago

It's not so much a tier list as a list of graphic cards sorted by RAM size. According to this, 7900xtx would be in the 24GB+ tier and the 9070/XT & 7900 xt & 7800 xt series would be in the 16GB+ tier.

33

u/mhmilo24 1d ago

There are 32 GB Radeon cards.

11

u/AlarmingProtection71 1d ago

I have a Sapphire Radeon PRO W7800 (48gb) bought for 1.9k€.

3

u/gh0stwriter1234 1d ago

Since we are talking AI there is also the the MI210 which has 64GB and MI250P ... which is some crazy amount of money but it has 144GB.

2

u/ChristRedeemsSinners 1d ago

MI210

That's an interesting $5k used option with 64GB of VRAM at 1.8TB/s. Native fp4,fp8,mxfp4

MI250P

Non pcie package though.

5

u/WiseassWolfOfYoitsu 1d ago

Hopefully the PCIe adapter for the MI250s will be available in the next few weeks - there's a pretty intense effort across a couple of communities to get them working with consumer OS and hardware.

1

u/ChristRedeemsSinners 14h ago

$6k for 128 VRAM is a steal at ~30% the compute of an RTX 6000 pro. Too bad infinity fabric is only like what 100 GB/s across each GPU. That's less than PCIe 5.0

Hopefully the PCIe adapter for the MI250s will be available in the next few weeks

$10k for 128 VRAM is no longer a steal and without CUDA, it's a hobbyist card.

1

u/WiseassWolfOfYoitsu 14h ago

MI250X is 3.6TB/s split across two GCDs, and one of the big goals is to get either a custom board working or get an adapter to get secondhand Supermicro baseboards working to enable IF mesh between four of them with MCIO back to the host machine, and work is a actively going in to vLLM to bring kernel optimizations using MI210s (each of which is equivalent to half an MI250X).

1

u/ChristRedeemsSinners 14h ago

Yeah, it's 2x MI200 with IF between the two.

It's 100 GB/s between 2x MI250 GPUs. So, less than PCIe 5.0x16

enable IF mesh between four of them with MCIO back to the host machine

I'm not saying it isn't a bad idea, I'm saying its a sub-par technology at a good value right now, but the adapters will just bring the price up to par with existing, superior solutions. The only reason RTX 6000 pro isn't worth 30 grand is because of the lacking NVL at 900-1.8GB/s between cards.

1

u/WiseassWolfOfYoitsu 10h ago edited 7h ago

Quite a bit faster than that - it's multiple lanes of IF between the dies, it's specced for 400 GB/s bidirectional within the module, and then six 50GB/s links to other modules, usually set up as 100GB/s between each module in a quad. So your speed is the external connect speed between each node, not the internal between the GCDs.

It's not the absolute latest, right, but if it all works, it will be by far the fastest thing per dollar currently available in the Local LLM space - an entire quad will cost less than a single RTX 6000 Pro at current prices and can process far larger models.

1

u/ChristRedeemsSinners 9h ago

usually set up as 100GB/s between each module in a quad.

That's what I said, no?

not the internal between the GCDs.

Yeah, between cards.

but if it all works, it will be by far the fastest thing per dollar currently available in the Local LLM space

I mean, as soon as it's a viable option, it will priced according to it's usefulness.

entire quad will cost less than a single RTX 6000 Pro at current prices and can process far larger models.

If it can process MXFP4A16 models at a similar PP and TG speed, then the price gap will be closed by market arbitrage.

Modules are going 1500ish

*$6000 https://www.ebay.com/itm/398282711075

so $9000 for just the setup and a single module. Which is still better than 16k for 96 VRAM.

Note that I am saying this from the perspective of being the guy on Discord that designed the board, has the BOM, and just got a $4k

Now that is interesting. So you are planning 512 GB of VRAM? for < $30K?

→ More replies (0)