IKBO handles broadcast inside kernels for recommendation model inference, avoiding materialized tensors and reducing memory movement. It reduces compute-intensive net latency by up to 2/3 and improves attention throughput by up to 2.4× and 6.4×. https://bit.ly/4umfoYY #PyTorch #OpenSourceAI
1 likes 0 replies
?