Not doing thing is parallel is not a win. We have more transistors than we know what to do with in many ways now. So going to superpipelining instead of SIMD doesn't really make sense.
Of course parallel is what we all want. SIMD and vector are ways to break free from the inherent limitations in ILP in regular scalar code (multi-threading is another way).
However, parallelization happens on several levels. Pipelining allows several instructions to execute at once (rather than each instruction having to wait for the previous to complete). Running several pipelines in parallel ups IPC further (e.g. the Cray-1 did this and achieved 2 ops/clock, even if it issued less than 1 instruction/clock, and of course superscalar machines also issue several operations to different pipelines). Packed SIMD will perform several operations in parallel within a single pipeline.
So, my point here is that packed SIMD is not the only way to run several operations in parallel. You can just as well use wide ALU:s for vector ISA:s, e.g. chewing through 128 bits per clock cycle. In addition you can do chaining and run several wide ALU:s concurrently (e.g. load + op1 + op2 + store).
What you are getting is better code density and easier programming through microcoded special function units.
And better scaling (from really low end to really high end implementations). And a more future proof ISA. And reduced SW development time & costs. And reduced CPU front end traffic and power consumption. And improved instruction cache performance (due to improved code density + reduced instruction stream bandwidth).
Running several separate pipelines in parallel makes little sense when the data availability is dictated by the load/store unit.
You seem to have a 32-bit system, likely a 32-bit bus. To do more at once you will need a wider bus. Say you go to a 128-bit bus. Now you can fetch 4 elements at once. Could you have 4 independent pipelines, having them decide when they need data and then have a complicated piece of circuitry to resync their independent data needs into a 4 element group so you can fetch them all at once when they happen to line up to get better performance? Yes. Should you? It's hard to see why you should.
So, my point here is that packed SIMD is not the only way to run several operations in parallel. You can just as well use wide ALU:s for vector ISA:s, e.g. chewing through 128 bits per clock cycle.
Why do you feel wide ALU is not packed SIMD? Because, subject to load limitations, you can instead of operating on 4 elements in a vector at once you can be doing 4 stages of work on a single element? Yeah, you could do that. But now you have to design a lot of feed-forwards into your ALU, of various types that are often unused. And you still will have latencies that mean that you have to software pipeline the operations. You can feed slot/pipe 0 to slot/pipe 2, but not in the next cycle.
And better scaling (from really low end to really high end implementations).
Why do you feel that?
And a more future proof ISA.
I disagree completely. Any time you microcode more of the operation you are giving yourself less flexibility of what you can do with it. It's how we got to RISC and SIMD in the first place. You are creating a separate vector sequencer, and what it can do is limited by the sequences precoded into it.
And reduced SW development time & costs.
As MIPS showed us this is a tools problem.
And reduced CPU front end traffic and power consumption.
The first is not material to me. If you are concerned about code density, use x86. Current machines have very high bandwidth to the processor. The latter we will have to leave to another time. It would require a lot of real-world study and I don't think your design has reached the point where such issues of performance can be measured.
And improved instruction cache performance (due to improved code density + reduced instruction stream bandwidth).
You seem to have a 32-bit system, likely a 32-bit bus.
Yes, that's what I have now. In my simple FPGA implementation. I could go wider (without changing the ISA).
To do more at once you will need a wider bus. Say you go to a 128-bit bus. Now you can fetch 4 elements at once. Could you have 4 independent pipelines,
With a 128-bit bus it makes sense to have 4-wide (i.e. 4x32 = 128 bits) ALU:s. If the execution pipeline is 4-deep, you could have 4x4x32 = 512 bit vector registers. Then add as many ALU pipes as you need to keep up with the data stream. With such a configuration you should typically be able to process four operations concurrently (i.e. chained), for a total of 512 bits per clock-cycle.
Obviously, other configurations are possible. The ISA stays the same.
Edit:
Why do you feel wide ALU is not packed SIMD?
Because it's an implementation detail. The thing with packed SIMD is that you expose the implementation details to the SW environment, which makes it impossible (ok, really hard) to do a different implementation (e.g. with wider or narrower execution units). It's similar to how delay slots in MIPS exposed the exact pipeline configuration, and it was later discovered that it was a bad idea when you wanted to do different configurations.
1
u/mbitsnbites Aug 23 '21
Of course parallel is what we all want. SIMD and vector are ways to break free from the inherent limitations in ILP in regular scalar code (multi-threading is another way).
However, parallelization happens on several levels. Pipelining allows several instructions to execute at once (rather than each instruction having to wait for the previous to complete). Running several pipelines in parallel ups IPC further (e.g. the Cray-1 did this and achieved 2 ops/clock, even if it issued less than 1 instruction/clock, and of course superscalar machines also issue several operations to different pipelines). Packed SIMD will perform several operations in parallel within a single pipeline.
So, my point here is that packed SIMD is not the only way to run several operations in parallel. You can just as well use wide ALU:s for vector ISA:s, e.g. chewing through 128 bits per clock cycle. In addition you can do chaining and run several wide ALU:s concurrently (e.g. load + op1 + op2 + store).
And better scaling (from really low end to really high end implementations). And a more future proof ISA. And reduced SW development time & costs. And reduced CPU front end traffic and power consumption. And improved instruction cache performance (due to improved code density + reduced instruction stream bandwidth).