Accelerate your transformer model with the new Block-Sparse-Flash-Attention! github.com/Danielohayon/B…
This training-free, drop-in replacement extends FlashAttention-2 with minimal code changes (CUDA Kernels Included). Paper: arxiv.org/abs/2512.07011
Associate Professor at the Technion, try to understand how AI works, and how to make it more efficient.
Haifa, Israel
Joined December 2019

