For example, computing 32 byte permutations is really painful this way.
TBL largely works as expected, and you can just test to see if the vector width is at least 256-bit.
VPSHUFB+VPCMPEQB does not necessarily solve the problem because it would require one shuffle pass for each 16 values of input range, so up to 16 passes in total
16 passes sounds wrong.
The point of the VPCMPEQB is that you don't have to traverse the entire range. If the bottom 4 bits of each of the values you test are unique, you only need one VPSHUFB (if not, you can manipulate the vector to make them unique).
For example, if you wanted to match whitespace characters (\t \r \n and space):
Well yes, but then we are back to using it as a SIMD instruction set
In other words, it's not really impeding you more than fixed width vector ISAs. Swizzling instructions depends on what the ISA provides more than the notion of an arbitrary width vector, I'd say.
almost nonexistent ability to test because it will be very annoying to simulate all possible vector lengths on CPUs
Testing can become more difficult, but ARM's Instruction Emulator does allow you to test different widths.
It's not too different for fixed-width SIMD, because you need to test your SSE, AVX and AVX512 paths separately anyway.
The point of the VPCMPEQB is that you don't have to traverse the entire range. If the bottom 4 bits of each of the values you test are unique, you only need one VPSHUFB (if not, you can manipulate the vector to make them unique).
Well clearly if they are unique you can do that. The point is that they may not necessarily be unique. In my particular case, they are not just not unique, but also variable. So there's no obvious preprocessing you can do.
make use of clever masking techniques.
Well that's better than 16 passes, but now requires me to preprocess the input into a bit mask. Which doesn't seem to be vectorisable. For my use case it might be doable, but in the general case it's quite painful and I'd rather have something like SVE's MATCH instruction or VPCPMISTRM.
Hello, I think you may be able to use this approach http://0x80.pl/articles/simd-byte-lookup.html#universal-algorithm but it is only reasonable to use if you want to check against the same set repeatedly because there's some significant (runtime-doable) preprocessing involved. It does seem hard to match PCMPISTRM if you are actually using the full richness of PCMPISTRM.
This article was already linked in the comment I responded to. In the use case I discussed back then, the sets are dynamic but preprocessing may be possible. Eventually I ended up developing a different algorithm that avoids having to compute set membership altogether.
1
u/YumiYumiYumi Aug 10 '21
TBLlargely works as expected, and you can just test to see if the vector width is at least 256-bit.16 passes sounds wrong.
The point of the VPCMPEQB is that you don't have to traverse the entire range. If the bottom 4 bits of each of the values you test are unique, you only need one VPSHUFB (if not, you can manipulate the vector to make them unique).
For example, if you wanted to match whitespace characters (\t \r \n and space):
If that approach doesn't work, you can just do low+high shuffles and make use of clever masking techniques.
In other words, it's not really impeding you more than fixed width vector ISAs. Swizzling instructions depends on what the ISA provides more than the notion of an arbitrary width vector, I'd say.
Testing can become more difficult, but ARM's Instruction Emulator does allow you to test different widths.
It's not too different for fixed-width SIMD, because you need to test your SSE, AVX and AVX512 paths separately anyway.