Ragnar {Groot Koerkamp} @curiouscoding.nl · Mar 13

In v2, we filp the SIMD tiling and bitpacking direction: Each lane represents a separate *pattern*, and the bits are adjacent *rows* of the matrix. Basically the original 'pack multiple patterns in a word' of Hyyrö 2005, but on an AVX-512 SIMD level. So now we can search 16 32bp patterns at once!

1 likes 1 replies

?

Replies

Ragnar {Groot Koerkamp} · Mar 13

Another cool trick, also by Hyyrö: if k=3 edits are allowed, then most of the time, this cost is exceeded after aligning around 6-10 characters. Thus, we first only find cost <=k candidate matches of the 16bp suffix of each pattern, so that we can double the SIMD parallellism to 32 x 16bp.