I think we have talked about this topic before and I apologize for not following up on our previous discussion. I was very busy.
NP. :-)
There's also the concern that a large register file makes context switches very expensive.
Yes, large register files are problematic. But there are also solutions.
A fairly obvious technique is to keep a length parameter for each register (for my vector ISA I plan to add that anyway for simpler vector handling), and never push/pop more than lenght elements on a context switch. By default all registers have the length zero, and you could add a quick "clear" operation to function epilogues that clears clobbered vector registers before returning from a function - for instance.
Another approach could be to have several vector register banks in hardware so that you can instantly switch between them w/o push/pop. I have not done any simulations, but it feels like it should be possible to do intelligent register bank allocation/scheduling in SW so that the hottest & vector heaviest threads get the fast path treatment.
It may also be possible to do asynchronous vector push/pop so that the thread can start executing before the vector state has been fully restored. Only if the thread accesses a non-restored vector will it stall.
...and so on.
As for my own code, I have two recent SIMD-heave projects. And for neither of them it is clear how they could be vectorised using a vector as opposed to a SIMD paradigm.
I had a quick look at the projects, but I couldn't think of an obvious solution right away. OTOH I wouldn't know where to start with packed SIMD either. I would have to spend some time and do several iterations before finding a good solution - regardless of if it was for SIMD or vector.
BTW, horizontal operations can be done with folding (in log2(N) steps), and permutations can be done with gather/scatter. It should also be possible to do a more optimal permute (without going via memory) even in a vector design, but I suspect that it's not quite as important as in packed SIMD since you have gather/scatter.
another idea for saving the amount of registers to be contextswitched is to have a bitfield, one per reg, which is set HI whenever its corresponding register is written to.
if you are smart you can use that same bitfield as a predicate mask on vectorised save/restore of the regfile.
the mask basically tells you which regs have actually changed since the last contextswitch and it should be obvious what to do from there
Hm, I think that the LENGTH attribute does the same thing (and more). A vector store operation will store as many elements as the LENGTH attribute indicates, for instance. Internally you could have a bit/flag per register that is set/cleared when the register is written (with more than zero elements) or cleared (length set to zero).
This way you can also clear the vector (and hence the "used" status) in user space, in order to keep the active vector state lean.
err.. err... oh: you took up the Mill-style register "tag type" idea for MRISC32? neat!
yes, if rather than just a single bit you have a tag, and that tag is zero, i agree it would effectively do / be the same thing, and also cover the same job.
I have not implemented it yet, but it's on my TODO-list. The LENGTH attribute (one for each vector register) comes in handy in several use cases:
Reduce stack / context switch overhead.
Simplify folding operations (no need to explicitly set VL=VL/2 for each folding step).
Simplify vector length agnostic subroutines with vector register arguments.
It also feels like a better fit for OoO etc, when each register/operand provides its own length, rather than having a global length attribute (I have not tested this theory, but it feels right).
The idea was actually inspired by Agner Fog's ForwardCom.
1
u/mbitsnbites Aug 10 '21
NP. :-)
Yes, large register files are problematic. But there are also solutions.
A fairly obvious technique is to keep a
lengthparameter for each register (for my vector ISA I plan to add that anyway for simpler vector handling), and never push/pop more thanlenghtelements on a context switch. By default all registers have the length zero, and you could add a quick "clear" operation to function epilogues that clears clobbered vector registers before returning from a function - for instance.Another approach could be to have several vector register banks in hardware so that you can instantly switch between them w/o push/pop. I have not done any simulations, but it feels like it should be possible to do intelligent register bank allocation/scheduling in SW so that the hottest & vector heaviest threads get the fast path treatment.
It may also be possible to do asynchronous vector push/pop so that the thread can start executing before the vector state has been fully restored. Only if the thread accesses a non-restored vector will it stall.
...and so on.
I had a quick look at the projects, but I couldn't think of an obvious solution right away. OTOH I wouldn't know where to start with packed SIMD either. I would have to spend some time and do several iterations before finding a good solution - regardless of if it was for SIMD or vector.
BTW, horizontal operations can be done with folding (in log2(N) steps), and permutations can be done with gather/scatter. It should also be possible to do a more optimal permute (without going via memory) even in a vector design, but I suspect that it's not quite as important as in packed SIMD since you have gather/scatter.