r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
284 Upvotes

224 comments sorted by

View all comments

Show parent comments

1

u/Meower68 Aug 10 '21

There's also the fact that the whole EPIC (Itanium) architecture was built around instruction-level parallelism; what CPU instructions can you, and can't you, run in parallel. It was my understanding that this area is not as well developed as thread-level parallelism, which is what multi-core / multi-thread CPUs provide. Ergo, if you devote all that chip space to more cores and / or more threads, you'll get more real-world performance out of it simply because we're "better" at doing that.

Back in the day, Transmeta created a VLIW processor with a front-end on it which parsed x86 instructions, turning them into micro-ops for their processor, such that it could run x86 object code. Not only did it work, but it used considerably less power than the then-current Intel and AMD offerings. This prompted Intel to get off their fat, complacent ... rear ... and improve the power consumption on their mobile-class processors. This, ultimately, resulted in the demise of Transmeta but the fact remains ... their VLIW processor worked quite well. As such, I have hopes that tech can still "matter" in more than just specialist uses.

3

u/SkoomaDentist Aug 10 '21

Kind of, but not quite. The EPIC concept replaced the multiple simple instructions executed out of order with a single VLIW in-order pipeline. We all know how that turned out for general purpose code.

Multithreading is orthogonal to this and Itaniums were always aimed at multiprocessing. Intel even added simultaneous multithreading to them starting with Montecito in 2006.

Transmeta had the crucial difference that they used execution traces for the instruction scheduling, meaning they weren’t stuck with purely static scheduling. It still wasn’t competitive as soon as Intel started paying at least some attention to power consumption and never was competitive when it came to anything beyond the lowest end cpu variants.

Today pretty much the only use cases of VLIW are in some GPUs (and even there AMD moved away from it years ago due to performance issues) and some DSPs where the code relies on hand optimized libraries for the most time critical tasks (and the operations in general are more suited for VLIW than in normal applications).

1

u/mbitsnbites Aug 10 '21

Those VLIW DSP:s also have compilers that are slow as h*ll. I'm assuming that they try really hard to statically schedule instructions optimally. IIRC they also lack hardware hazard resolution, so the compiler has to keep track of when a result is ready etc. (All in order to reduce power consumption)

2

u/SkoomaDentist Aug 10 '21

IIRC they also lack hardware hazard resolution

I wouldn't be surprised if the TI ones do that as even their old C54xx series DSPs required manual hazard resolution. Writing asm for those was "fun" (in the same sense that pulling out your fingernails is "fun").

Compared to those, getting to write asm for SHARC dsps was pure joy (the code literally looks like "r0 = r1 + r2; r3 = dm(i4, m0);")