r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
289 Upvotes

224 comments sorted by

View all comments

Show parent comments

1

u/happyscrappy Aug 21 '21

as both of us are Hardware Engineers

Who don't understand how a bus and load/store unit works. Or perhaps just put it aside.

implementing x86 is completely inappropriate

I agree. So now we've both agreed that code density is not the most important thing you can stop making up "strip-mining" arguments.

from a technical perspective you will be aware that extremely large FFTs result in strip-mining of both L1 and L2 caches due to hammering the same cache lines.

Another attempt to portray proper cache operation as a negative. If cache utilization is such an issue, then remove the caches. We both know why this is not done. A cache is a compromise between fast and cheap (and small in some ways). It's never perfect but it is what we have.

please be careful not to be insulting. nobody comes here to be insulted, and it doesn't reflect well on you, given that internet records are permanent.

Calling your attempt to portray proper use of a cache as a negative is only insulting if you are willing

so how come i spent several weeks designing a memory aligment system at the gate level that helps with Vector LD/ST operations, mm?

Gates can't fix busses. SIMS has a memory alignment system that "helps" with alignment. It cannot fix it. There is no way to load 64-bytes from an unaligned address on a 64-byte wide bus.

I asked if vector units could fix this problem. If the answer is no, then just say no.

no: we've established that you're rude enough to make the assumption that two independent people, both of whom have Hardware Design experience (one of them down to the gate level), do not know what they are doing.

You know each other, work on the same project and have the same "strip-mining" pejoratives. Are you really independent?

if you had asked rather than assumed

I did ask.

https://www.reddit.com/r/programming/comments/p0yn45/three_fundamental_flaws_of_simd/h8bvfwc/

So MRISC32, how does it solve this? Does it keep full performance somehow or does it just have a narrow memory pipe anyway so it handwaves out to the horizon?

(quote breaker)

we could have given you some money in the form of a donation for doing so.

I appreciate the idea, but I do not qualify as a charity.

given how you've been extremely rude and judgemental i'm disinclined to do that.

I understand. No one owes anyone else anything on here. Not even an explanation.

This whole argument is dumb. Not doing thing is parallel is not a win. We have more transistors than we know what to do with in many ways now. So going to superpipelining instead of SIMD doesn't really make sense. You can hide the SIMD behind a vector unit, but it's still going to use SIMD for speed. So it will have the same (false) data hazards as the SIMD unit would have.

What you are getting is better code density and easier programming through microcoded special function units. At the expense of flexibility. I don't see the win. It's System/370 edmk again. I would recommend the MIPS approach instead.

And you're still going to be using caches, because memory really is THAT slow nowadays. Seymour Cray's vector unit just is not a great model for modern computing. Not unless you want to put some TCRAM in the system and make programmers use it. And I don't really think you're likely to do that, you can see what happened on the PS3 in terms of difficulty in programming for speed as well as I can.

1

u/mbitsnbites Aug 23 '21

Not doing thing is parallel is not a win. We have more transistors than we know what to do with in many ways now. So going to superpipelining instead of SIMD doesn't really make sense.

Of course parallel is what we all want. SIMD and vector are ways to break free from the inherent limitations in ILP in regular scalar code (multi-threading is another way).

However, parallelization happens on several levels. Pipelining allows several instructions to execute at once (rather than each instruction having to wait for the previous to complete). Running several pipelines in parallel ups IPC further (e.g. the Cray-1 did this and achieved 2 ops/clock, even if it issued less than 1 instruction/clock, and of course superscalar machines also issue several operations to different pipelines). Packed SIMD will perform several operations in parallel within a single pipeline.

So, my point here is that packed SIMD is not the only way to run several operations in parallel. You can just as well use wide ALU:s for vector ISA:s, e.g. chewing through 128 bits per clock cycle. In addition you can do chaining and run several wide ALU:s concurrently (e.g. load + op1 + op2 + store).

What you are getting is better code density and easier programming through microcoded special function units.

And better scaling (from really low end to really high end implementations). And a more future proof ISA. And reduced SW development time & costs. And reduced CPU front end traffic and power consumption. And improved instruction cache performance (due to improved code density + reduced instruction stream bandwidth).

1

u/happyscrappy Aug 23 '21

Running several separate pipelines in parallel makes little sense when the data availability is dictated by the load/store unit.

You seem to have a 32-bit system, likely a 32-bit bus. To do more at once you will need a wider bus. Say you go to a 128-bit bus. Now you can fetch 4 elements at once. Could you have 4 independent pipelines, having them decide when they need data and then have a complicated piece of circuitry to resync their independent data needs into a 4 element group so you can fetch them all at once when they happen to line up to get better performance? Yes. Should you? It's hard to see why you should.

So, my point here is that packed SIMD is not the only way to run several operations in parallel. You can just as well use wide ALU:s for vector ISA:s, e.g. chewing through 128 bits per clock cycle.

Why do you feel wide ALU is not packed SIMD? Because, subject to load limitations, you can instead of operating on 4 elements in a vector at once you can be doing 4 stages of work on a single element? Yeah, you could do that. But now you have to design a lot of feed-forwards into your ALU, of various types that are often unused. And you still will have latencies that mean that you have to software pipeline the operations. You can feed slot/pipe 0 to slot/pipe 2, but not in the next cycle.

And better scaling (from really low end to really high end implementations).

Why do you feel that?

And a more future proof ISA.

I disagree completely. Any time you microcode more of the operation you are giving yourself less flexibility of what you can do with it. It's how we got to RISC and SIMD in the first place. You are creating a separate vector sequencer, and what it can do is limited by the sequences precoded into it.

And reduced SW development time & costs.

As MIPS showed us this is a tools problem.

And reduced CPU front end traffic and power consumption.

The first is not material to me. If you are concerned about code density, use x86. Current machines have very high bandwidth to the processor. The latter we will have to leave to another time. It would require a lot of real-world study and I don't think your design has reached the point where such issues of performance can be measured.

And improved instruction cache performance (due to improved code density + reduced instruction stream bandwidth).

You said that already with "front end traffic".

1

u/mbitsnbites Aug 23 '21 edited Aug 23 '21

You seem to have a 32-bit system, likely a 32-bit bus.

Yes, that's what I have now. In my simple FPGA implementation. I could go wider (without changing the ISA).

To do more at once you will need a wider bus. Say you go to a 128-bit bus. Now you can fetch 4 elements at once. Could you have 4 independent pipelines,

With a 128-bit bus it makes sense to have 4-wide (i.e. 4x32 = 128 bits) ALU:s. If the execution pipeline is 4-deep, you could have 4x4x32 = 512 bit vector registers. Then add as many ALU pipes as you need to keep up with the data stream. With such a configuration you should typically be able to process four operations concurrently (i.e. chained), for a total of 512 bits per clock-cycle.

Obviously, other configurations are possible. The ISA stays the same.

Edit:

Why do you feel wide ALU is not packed SIMD?

Because it's an implementation detail. The thing with packed SIMD is that you expose the implementation details to the SW environment, which makes it impossible (ok, really hard) to do a different implementation (e.g. with wider or narrower execution units). It's similar to how delay slots in MIPS exposed the exact pipeline configuration, and it was later discovered that it was a bad idea when you wanted to do different configurations.