🔗 Our new blog looks at how FP4 is moving beyond compression into a practical primitive for training and inference across both LLMs and diffusion models:
research.nvidia.com/labs/eai/blogs…
1. Why Four Bits Is Hard: Only 15 values make scaling critical.
2. NVFP4: Smaller blocks and finer
NVIDIA Research Intern @NVIDIA
The University of Hong Kong PhD ing @HKUniversity.
Efficient AI; Long AI; LLM/AIGC; Embodied
Hong Kong


