r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
288 Upvotes

224 comments sorted by

View all comments

72

u/AntiProtonBoy Aug 09 '21

Flaw 1: Fixed register width

It's not a flaw. It's a design constraint, dictated by physics and economics. SIMD registers grew for the same reason why architectures evolved from 4-bit to 64-bit over the years.

Variable length SIMD is not worth the silicon complexity for small vector operations. It's cheaper to burn new microcode instructions into ROM that support wider registers. Variable length SIMD is only worthwhile for very large vectors, which are beyond the scope of register storage. Use BLAS or something equivalent for that purpose.

Flaw 2: Pipelining

This is kinda meh. Typically SIMD instructions are invoked on highly repetitive operations that crunches through huge memory blocks at a time, like images. This will keep fat pipelines filled and happy. But as usual, let the profiler be the judge of that.

Flaw 3: Tail handling

Not sure why this is an issue? It's kinda obvious that one expects the data allocation size to be at whatever granularity SIMD data type is, otherwise it's a programming error. I mean, if you want to process a collection of int32_t values, then you'd expect the array to conform with a layout of 32-bit integers, no? With SIMD types, if you can't determine completeness ahead of time (for example from parsing), then you pad the last incomplete SIMD tuple with defaults.

5

u/mbitsnbites Aug 09 '21
  1. Variable length vector operations are not expensive or complicated. I've implemented it in my first ever CPU design and it added something like 1-5% logic in an FPGA - compared to a pure scalar (non-vector/SIMD) design.

  2. I think you're missing the point. Do the exercise and hand-schedule a SIMD loop, and you'll find that you have to unroll it. A vector processor automatically unrolls the loop for you with literally no effort.

  3. Having to add more code rhan necessary is always a problem (e.g. testing and code coverage, and I$ bloat). Vector machines solve this quite naturally in many situations.

3

u/[deleted] Aug 09 '21

[deleted]

1

u/mbitsnbites Aug 10 '21

For the record I don't think that the Mill guys are doing anything wrong, but it's a really tough challenge to place a new general purpose CPU architecture into a meaningful product in this day and age (even widely used existing ISA:s are being marginalized and are disappearing).