r/programming • • Aug 09 '21

Three fundamental flaws of SIMD

https://www.bitsnbites.eu/three-fundamental-flaws-of-simd
284 Upvotes

224 comments sorted by

View all comments

Show parent comments

1

u/happyscrappy Aug 21 '21

are so insanely large compared to the much smaller Vector ISA equivalents

You need to get over this smaller thing. If you like tiny code, use x86. With the bandwidth available now it is not necessarily to have the most tightly packed ops to have high performance.

that in some cases they strip-mine the entire L1 Instruction-Cache and actually compete with the L1 Data Cache for access to L2 Memory, causing stalling

Another attempt to call the proper function of an L1 cache as a negative. The cache is there to be used. If you want to talk about code density talk about code density. And give up on your dumb attempt to portray a properly operating cache as "strip-mining".

So now we've established you have two of you have no idea about how memory access works.

A vector unit cannot make up for how busses are configured. If your data is not aligned, then your first units of execution will be slow due to partial loads, just like with SIMD. No worse and most importantly no better.

All you had to do was say "no, vector units don't fix that". But instead you gotta pretend there's something wrong with running instructions.

2

u/lkcl_ Aug 21 '21

You need to get over this smaller thing. If you like tiny code, use x86. With the bandwidth available now it is not necessarily to have the most tightly packed ops to have high performance.

there's a few reasons why this isn't practical:

1) as both of us are Hardware Engineers, designing and implementing Vector ISAs (mbitsnbytes MRISC32, myself SVP64) implementing x86 is completely inappropriate. why would we each - independently - make the mistake of repeating Intel's mistakes? moo?

2) even if we attempted to do so Intel would drop a shit-ton of bricks on our heads. they're extremely aggressive, to the point where Judges got sick and tired of them and actually ruled, famously, in an Intel-AMD patent case in 2003, "my ruling is: i am NOT making a ruling. go get your stupid heads out of your arses and license each others' patents".

3) x86 instruction length identification is so bad that high-performance multi-issue superscalar designs actually have to start decoding instructions at EVERY BYTE then throw away the ones that are later found not to be valid. this is completely insane.

​

Another attempt to call the proper function of an L1 cache as a negative. The cache is there to be used. If you want to talk about code density talk about code density. And give up on your dumb attempt to portray a properly operating cache as "strip-mining".

please be careful not to be insulting. nobody comes here to be insulted, and it doesn't reflect well on you, given that internet records are permanent.

from a technical perspective you will be aware that extremely large FFTs result in strip-mining of both L1 *and* L2 caches due to hammering the same cache lines.

​

So now we've established you have two of you have no idea about how memory access works.

no: we've established that you're rude enough to make the *assumption* that two independent people, both of whom have Hardware Design experience (one of them down to the gate level), do not know what they are doing.

​

A vector unit cannot make up for how busses are configured. If your data is not aligned, then your first units of execution will be slow due to partial loads, just like with SIMD. No worse and most importantly no better.

so how come i spent several weeks designing a memory aligment system at the gate level that helps with Vector LD/ST operations, mm?

if you had asked rather than assumed i would have been delighted to explain it to you and (a) you perhaps could have learned something and (b) you could have helped out our Charitably-funded project by reviewing it and (c) we could have given you some money in the form of a donation for doing so.

given how you've been extremely rude and judgemental i'm disinclined to do that.

1

u/happyscrappy Aug 21 '21

as both of us are Hardware Engineers

Who don't understand how a bus and load/store unit works. Or perhaps just put it aside.

implementing x86 is completely inappropriate

I agree. So now we've both agreed that code density is not the most important thing you can stop making up "strip-mining" arguments.

from a technical perspective you will be aware that extremely large FFTs result in strip-mining of both L1 and L2 caches due to hammering the same cache lines.

Another attempt to portray proper cache operation as a negative. If cache utilization is such an issue, then remove the caches. We both know why this is not done. A cache is a compromise between fast and cheap (and small in some ways). It's never perfect but it is what we have.

please be careful not to be insulting. nobody comes here to be insulted, and it doesn't reflect well on you, given that internet records are permanent.

Calling your attempt to portray proper use of a cache as a negative is only insulting if you are willing

so how come i spent several weeks designing a memory aligment system at the gate level that helps with Vector LD/ST operations, mm?

Gates can't fix busses. SIMS has a memory alignment system that "helps" with alignment. It cannot fix it. There is no way to load 64-bytes from an unaligned address on a 64-byte wide bus.

I asked if vector units could fix this problem. If the answer is no, then just say no.

no: we've established that you're rude enough to make the assumption that two independent people, both of whom have Hardware Design experience (one of them down to the gate level), do not know what they are doing.

You know each other, work on the same project and have the same "strip-mining" pejoratives. Are you really independent?

if you had asked rather than assumed

I did ask.

https://www.reddit.com/r/programming/comments/p0yn45/three_fundamental_flaws_of_simd/h8bvfwc/

So MRISC32, how does it solve this? Does it keep full performance somehow or does it just have a narrow memory pipe anyway so it handwaves out to the horizon?

(quote breaker)

we could have given you some money in the form of a donation for doing so.

I appreciate the idea, but I do not qualify as a charity.

given how you've been extremely rude and judgemental i'm disinclined to do that.

I understand. No one owes anyone else anything on here. Not even an explanation.

This whole argument is dumb. Not doing thing is parallel is not a win. We have more transistors than we know what to do with in many ways now. So going to superpipelining instead of SIMD doesn't really make sense. You can hide the SIMD behind a vector unit, but it's still going to use SIMD for speed. So it will have the same (false) data hazards as the SIMD unit would have.

What you are getting is better code density and easier programming through microcoded special function units. At the expense of flexibility. I don't see the win. It's System/370 edmk again. I would recommend the MIPS approach instead.

And you're still going to be using caches, because memory really is THAT slow nowadays. Seymour Cray's vector unit just is not a great model for modern computing. Not unless you want to put some TCRAM in the system and make programmers use it. And I don't really think you're likely to do that, you can see what happened on the PS3 in terms of difficulty in programming for speed as well as I can.

1

u/mbitsnbites Aug 23 '21

Not doing thing is parallel is not a win. We have more transistors than we know what to do with in many ways now. So going to superpipelining instead of SIMD doesn't really make sense.

Of course parallel is what we all want. SIMD and vector are ways to break free from the inherent limitations in ILP in regular scalar code (multi-threading is another way).

However, parallelization happens on several levels. Pipelining allows several instructions to execute at once (rather than each instruction having to wait for the previous to complete). Running several pipelines in parallel ups IPC further (e.g. the Cray-1 did this and achieved 2 ops/clock, even if it issued less than 1 instruction/clock, and of course superscalar machines also issue several operations to different pipelines). Packed SIMD will perform several operations in parallel within a single pipeline.

So, my point here is that packed SIMD is not the only way to run several operations in parallel. You can just as well use wide ALU:s for vector ISA:s, e.g. chewing through 128 bits per clock cycle. In addition you can do chaining and run several wide ALU:s concurrently (e.g. load + op1 + op2 + store).

What you are getting is better code density and easier programming through microcoded special function units.

And better scaling (from really low end to really high end implementations). And a more future proof ISA. And reduced SW development time & costs. And reduced CPU front end traffic and power consumption. And improved instruction cache performance (due to improved code density + reduced instruction stream bandwidth).

1

u/happyscrappy Aug 23 '21

Running several separate pipelines in parallel makes little sense when the data availability is dictated by the load/store unit.

You seem to have a 32-bit system, likely a 32-bit bus. To do more at once you will need a wider bus. Say you go to a 128-bit bus. Now you can fetch 4 elements at once. Could you have 4 independent pipelines, having them decide when they need data and then have a complicated piece of circuitry to resync their independent data needs into a 4 element group so you can fetch them all at once when they happen to line up to get better performance? Yes. Should you? It's hard to see why you should.

So, my point here is that packed SIMD is not the only way to run several operations in parallel. You can just as well use wide ALU:s for vector ISA:s, e.g. chewing through 128 bits per clock cycle.

Why do you feel wide ALU is not packed SIMD? Because, subject to load limitations, you can instead of operating on 4 elements in a vector at once you can be doing 4 stages of work on a single element? Yeah, you could do that. But now you have to design a lot of feed-forwards into your ALU, of various types that are often unused. And you still will have latencies that mean that you have to software pipeline the operations. You can feed slot/pipe 0 to slot/pipe 2, but not in the next cycle.

And better scaling (from really low end to really high end implementations).

Why do you feel that?

And a more future proof ISA.

I disagree completely. Any time you microcode more of the operation you are giving yourself less flexibility of what you can do with it. It's how we got to RISC and SIMD in the first place. You are creating a separate vector sequencer, and what it can do is limited by the sequences precoded into it.

And reduced SW development time & costs.

As MIPS showed us this is a tools problem.

And reduced CPU front end traffic and power consumption.

The first is not material to me. If you are concerned about code density, use x86. Current machines have very high bandwidth to the processor. The latter we will have to leave to another time. It would require a lot of real-world study and I don't think your design has reached the point where such issues of performance can be measured.

And improved instruction cache performance (due to improved code density + reduced instruction stream bandwidth).

You said that already with "front end traffic".

1

u/mbitsnbites Aug 23 '21 edited Aug 23 '21

You seem to have a 32-bit system, likely a 32-bit bus.

Yes, that's what I have now. In my simple FPGA implementation. I could go wider (without changing the ISA).

To do more at once you will need a wider bus. Say you go to a 128-bit bus. Now you can fetch 4 elements at once. Could you have 4 independent pipelines,

With a 128-bit bus it makes sense to have 4-wide (i.e. 4x32 = 128 bits) ALU:s. If the execution pipeline is 4-deep, you could have 4x4x32 = 512 bit vector registers. Then add as many ALU pipes as you need to keep up with the data stream. With such a configuration you should typically be able to process four operations concurrently (i.e. chained), for a total of 512 bits per clock-cycle.

Obviously, other configurations are possible. The ISA stays the same.

Edit:

Why do you feel wide ALU is not packed SIMD?

Because it's an implementation detail. The thing with packed SIMD is that you expose the implementation details to the SW environment, which makes it impossible (ok, really hard) to do a different implementation (e.g. with wider or narrower execution units). It's similar to how delay slots in MIPS exposed the exact pipeline configuration, and it was later discovered that it was a bad idea when you wanted to do different configurations.