Got sick of hand-tuning GPU kernels, so we built a compiler.
Photon 2.0 compiles Moondream, Qwen 3.5, and Gemma 4 into megakernels: the entire forward pass in one GPU program.
I think this is one of the most fascinating projects I’ve tinkered on yet. Parameterizing our weights as a learned weighted blend of symexp and regular linear weights gives up to ~1.42x training speedup wallclock, and can be fused back into standard weights for deployment🧵