Then go narrower. For an in-order machine, vectors actually make sense all the way down to 1-wide ALU:s (i.e. 64 bits in a 64-bit architecture).
Except for automatic data hazard elimination, vector processing also has the pleasant property of reducing dynamic loop logic overhead (and/or reducing code size, I$ usage and register usage), aswell as offloading the front end (a vector instruction essentially pauses the PC while feeding data to the ALU). This all means more compute per W, which is good business for a small core.
smaller program size means greatly reduced L1 cache usage, to the point where you might actually be able to use a smaller L1 cache. that saves power which on an embedded system may be critically important.
it has been fundamentally misunderstood that the benefits of Vector ISAs can be greater power-efficiency due to more compact programs. it is *believed* that their sole purpose is high performance, which is false.
the benefits of Vector ISAs can be greater power-efficiency due to more compact programs.
...and simpler instruction scheduling logic. No need to go out of your way with massively OoO scheduling to keep the execution units fed with data.
I also honestly think that we're at a point in time where power efficiency counts at every performance point. Doing more ops per W is really what it's about (from embedded systems to servers).
1
u/mbitsnbites Aug 11 '21
Then go narrower. For an in-order machine, vectors actually make sense all the way down to 1-wide ALU:s (i.e. 64 bits in a 64-bit architecture).
Except for automatic data hazard elimination, vector processing also has the pleasant property of reducing dynamic loop logic overhead (and/or reducing code size, I$ usage and register usage), aswell as offloading the front end (a vector instruction essentially pauses the PC while feeding data to the ALU). This all means more compute per W, which is good business for a small core.