My current research focuses on the post-training of Large Language Models (LLMs) and Multimodal LLMs (MLLMs), spanning Reinforcement Learning (RL) and On-Policy Distillation (OPD). I am broadly interested in a range of topics including reasoning enhancement, long-horizon tasks, alignment of Models, and pushing the boundaries of (M)LLMs to unlock their full potential in complex real-world scenarios.