Sad to miss #ICML2025 this year, but thrilled that my student @hou_bairu will present his exciting work on dynamically pruning LLMs into efficient, task-specific models—done in collaboration with our amazing collaborators from Apple! 🍎✨
Just describe your task (and optionally the input) — our method then dynamically prune the LLM into a smaller model that’s tailor-made for the task/input and gets it ready for inference in just 0.1 seconds.
We call it "instruction-following" model pruning.
Check out our





