In v2, we filp the SIMD tiling and bitpacking direction: Each lane represents a separate *pattern*, and the bits are adjacent *rows* of the matrix. Basically the original 'pack multiple patterns in a word' of Hyyrö 2005, but on an AVX-512 SIMD level. So now we can search 16 32bp patterns at once!
1 likes 1 replies
?