Archit3chan hour ago
Hot take: there is no portable SIMD.
You can either have performance (=write manual ASM for each platform), or portability, but not both.
What so-called "portable SIMD" libraries give you is "portable auto-vectorization". "Portable performance" is a global property of the algorithm. Relying on auto-vectorization will result in e.g. sub-optimal register spills in practice. The microbenchmarks will look great, though. ;)
Sharlin30 minutes ago
Getting 2x or 4x performance in your inner loops using a reasonable SIMD library is infinitely better than theoretically getting 8x performance with hand-coded nonportable intrinsics, because the latter is never going to happen in most programs, so the actual point of comparison is scalar code, or autovectorized code at best.
louthy37 minutes ago
> there is no portable SIMD
Except in languages with a JIT compiler
Tanjreeve27 minutes ago
Starts to get a bit philosophical on what constitutes "portable" but JIT compilers would emit an opcode based off of whatever the frontend/IR is saying to do surely?
IshKebab15 minutes ago
It's a continuum. Some things basically all SIMD implementations support. Want to add 2 4xf32 vectors together? That's pretty easy to do portably.
But yeah to be fair if you are at that point, you probably want to go fully non-portable anyway. Especially with AI.
Has anyone even figured out how to do vector stuff (SVE/RVV) without assembly?
Scene_Cast235 minutes ago
What about numpy, numba, and torch.compile?
izacus3 minutes ago
Those are manually optimized per arch, aren't they?
justmeeewan hour ago
[flagged]