1. X
  2. Haoran He ✈️ ICML26
Log inSign up
Haoran He ✈️ ICML26
153 posts
Image
user avatar
Haoran He ✈️ ICML26
@tinner_he
Building multimodal (agentic) RL | Ph.D. student at @hkust, B.Eng. from @SJTU1896
Hong Kong
tinnerhrhe.github.io
Joined June 2019
404
Following
224
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Haoran He ✈️ ICML26
    @tinner_he
    Jan 6
    Reinforcement Learning (RL) is the key to aligning diffusion models, but it comes with a curse: Reward Hacking. 🎭 Models often game the proxy reward (e.g., OCR scores) while destroying image quality. ⚡ Introducing GARDO: Reinforcing Diffusion Models without Reward Hacking. 👇
    Image
  • user avatar
    Haoran He ✈️ ICML26
    @tinner_he
    Apr 23
    🧐Wonder to know how random policy valuation contributes to LLM reasoning? Come to check our poster at Pavilion 3 P3-#1712, 3:15 pm - 5:45 pm! #ICLR2026
    Image
    Image
    user avatar
    Haoran He ✈️ ICML26
    @tinner_he
    Sep 30, 2025
    🚨Our new paper: Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards We challenge the RL status quo. We find you don't need complex policy optimization for top-tier math reasoning. The key? Evaluating the Q function of a simple uniformly random policy. 🤯
  • user avatar
    Haoran He ✈️ ICML26
    @tinner_he
    Mar 25
    Science has no nationality. We hope @NeurIPSConf reconsider the sanction policy. #NeurIPS #AI
    Image
    Image
  • user avatar
    Haoran He ✈️ ICML26
    @tinner_he
    Feb 12
    🥳Come to submit your papers to our 1st Workshop on Video World Models! See you at CVPR 2026!
    user avatar
    Jiwen Yu
    @yujiwenHK
    Feb 12
    🚀 Excited to announce the 1st Workshop on Video World Models @ CVPR 2026! 🎓Organized by researchers from Stanford, Oxford, HKU, NUS & NVIDIA. 📝 Call for papers is open! 🔗 videoworldmodel-workshop.github.io
    Image
  • user avatar
    Haoran He ✈️ ICML26
    @tinner_he
    Jan 26
    🚀ROVER has been accepted by ICLR 2026! See you at 🇧🇷! paper: arxiv.org/abs/2509.24981
    user avatar
    Haoran He ✈️ ICML26
    @tinner_he
    Sep 30, 2025
    🚨Our new paper: Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards We challenge the RL status quo. We find you don't need complex policy optimization for top-tier math reasoning. The key? Evaluating the Q function of a simple uniformly random policy. 🤯
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement