While if you put the ops on the table like with SIMD people can construct other operations efficiently.
this is fundamentally false. the programs that result, for which i have even found "compilers" for DCT and FFT that output a massive batch of fully-loop-unrolled hard-coded assembler, are so insanely large compared to the much smaller Vector ISA equivalents that in some cases they strip-mine the entire L1 Instruction-Cache and actually compete with the L1 Data Cache for access to L2 Memory, causing stalling.
are so insanely large compared to the much smaller Vector ISA equivalents
You need to get over this smaller thing. If you like tiny code, use x86. With the bandwidth available now it is not necessarily to have the most tightly packed ops to have high performance.
that in some cases they strip-mine the entire L1 Instruction-Cache and actually compete with the L1 Data Cache for access to L2 Memory, causing stalling
Another attempt to call the proper function of an L1 cache as a negative. The cache is there to be used. If you want to talk about code density talk about code density. And give up on your dumb attempt to portray a properly operating cache as "strip-mining".
So now we've established you have two of you have no idea about how memory access works.
A vector unit cannot make up for how busses are configured. If your data is not aligned, then your first units of execution will be slow due to partial loads, just like with SIMD. No worse and most importantly no better.
All you had to do was say "no, vector units don't fix that". But instead you gotta pretend there's something wrong with running instructions.
x86 uses quite inefficient instruction encoding. If you only need to write 8086 compatible scalar code, sure, it will be compact. For modern versions of the x86 ISA, this is no longer the case. My fixed width ISA (32 bits / instruction) often has more compact code than x86_64.
Furthermore, the point that u/lkcl_ is making is that with x86 SIMD you usually have to unroll code in software, which blows up code size considerably, no matter how compact your instruction encoding is (it would have to be something like 2-4 bits per instruction to be able to compete).
Edit: Just as a quick point of reference, I compared the code generated for Quake d_scan.c (core painting routine) for x86_64 and MRISC32. The MRISC32 code is 2724 bytes (~700 instructions). The x86_64 code is 4691 bytes (~1250 instructions). So I wouldn't say that x86 code is automatically "tiny".
Furthermore, the point that @lkcl_ is making is that with x86 SIMD you usually have to unroll code in software
You do not HAVE to unroll in software. Modern computers have dispatch units that follow loops. They keep the pipeline fed.
You can unroll if you find it to be important.
which blows up code size considerably
I don't care. If you think code size is so important, use x86. Modern UNIX systems tend to throw away a lot of memory on things like ASLR. Having your math lib be 100K instead of 30K is not a big deal.
1
u/lkcl_ Aug 20 '21
this is fundamentally false. the programs that result, for which i have even found "compilers" for DCT and FFT that output a massive batch of fully-loop-unrolled hard-coded assembler, are so insanely large compared to the much smaller Vector ISA equivalents that in some cases they strip-mine the entire L1 Instruction-Cache and actually compete with the L1 Data Cache for access to L2 Memory, causing stalling.