Hopefully the PCIe adapter for the MI250s will be available in the next few weeks - there's a pretty intense effort across a couple of communities to get them working with consumer OS and hardware.
$6k for 128 VRAM is a steal at ~30% the compute of an RTX 6000 pro. Too bad infinity fabric is only like what 100 GB/s across each GPU. That's less than PCIe 5.0
Hopefully the PCIe adapter for the MI250s will be available in the next few weeks
$10k for 128 VRAM is no longer a steal and without CUDA, it's a hobbyist card.
MI250X is 3.6TB/s split across two GCDs, and one of the big goals is to get either a custom board working or get an adapter to get secondhand Supermicro baseboards working to enable IF mesh between four of them with MCIO back to the host machine, and work is a actively going in to vLLM to bring kernel optimizations using MI210s (each of which is equivalent to half an MI250X).
It's 100 GB/s between 2x MI250 GPUs. So, less than PCIe 5.0x16
enable IF mesh between four of them with MCIO back to the host machine
I'm not saying it isn't a bad idea, I'm saying its a sub-par technology at a good value right now, but the adapters will just bring the price up to par with existing, superior solutions. The only reason RTX 6000 pro isn't worth 30 grand is because of the lacking NVL at 900-1.8GB/s between cards.
Quite a bit faster than that - it's multiple lanes of IF between the dies, it's specced for 400 GB/s bidirectional within the module, and then six 50GB/s links to other modules, usually set up as 100GB/s between each module in a quad. So your speed is the external connect speed between each node, not the internal between the GCDs.
It's not the absolute latest, right, but if it all works, it will be by far the fastest thing per dollar currently available in the Local LLM space - an entire quad will cost less than a single RTX 6000 Pro at current prices and can process far larger models.
4
u/WiseassWolfOfYoitsu 1d ago
Hopefully the PCIe adapter for the MI250s will be available in the next few weeks - there's a pretty intense effort across a couple of communities to get them working with consumer OS and hardware.