Pinned
Hao AI Lab at UCSD. Our mission is to democratize large machine learning models, algorithms, and their underlying systems.
Joined March 2024
- (1/6) The FastVideo team is excited to release a new set of FP4 attention kernels for B300, achieving up to 1.69x speedup over FA4!
- Thrilled to share our poster sessions at #ICML2026 in Seoul! ☕️🇰🇷 Stop by to chat with our lab members and collaborators about parallel decoding, diffusion LLMs, speculative decoding, video sparse attention, quantization, and more. We’re excited to connect! 1️⃣ Strategic
- Introducing JetSpec: we find speculative decoding can push LLM generation latency to extreme by co-optimizing drafting cost and drafting quality with causal parallel tree drafting. JetSpec reaches up to 9.64x end-to-end speedup on MATH-500 and 4.58x on open-ended chat while



