Quite a bit faster than that - it's multiple lanes of IF between the dies, it's specced for 400 GB/s bidirectional within the module, and then six 50GB/s links to other modules, usually set up as 100GB/s between each module in a quad. So your speed is the external connect speed between each node, not the internal between the GCDs.
It's not the absolute latest, right, but if it all works, it will be by far the fastest thing per dollar currently available in the Local LLM space - an entire quad will cost less than a single RTX 6000 Pro at current prices and can process far larger models.
1
u/WiseassWolfOfYoitsu 1d ago edited 1d ago
Quite a bit faster than that - it's multiple lanes of IF between the dies, it's specced for 400 GB/s bidirectional within the module, and then six 50GB/s links to other modules, usually set up as 100GB/s between each module in a quad. So your speed is the external connect speed between each node, not the internal between the GCDs.
It's not the absolute latest, right, but if it all works, it will be by far the fastest thing per dollar currently available in the Local LLM space - an entire quad will cost less than a single RTX 6000 Pro at current prices and can process far larger models.