Gemma 4 with image inputs is very slow with the Triton attention backend because it doesn't prune the sliding window.
Small fix, 3-4x faster E2E inference 🫡
I like to shrink dem models 🤏
ML/AI Perf @dropbox
Prev. Co-Founder & Principal Scientist @mobius_labs (acquired by @dropbox)
PhD @inria



