Log inSign up
Skander Moalla
166 posts
Skander Moalla profile banner
@SkanderMoalla

Skander Moalla

@SkanderMoalla
RS Intern @ Meta FAIR | PhD @ EPFL, Caglar Gulcehre Lab for AI Research (CLAIRE) | Reinforcement learning, Large Language Model post-training
Lausanne, Switzerland
skandermoalla.com
Joined January 2017
459
Following
297
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @SkanderMoalla
    Skander Moalla
    @SkanderMoalla
    Sep 19, 2025
    See you at #NeurIPS2025! 🙌
    @SkanderMoalla
    Skander Moalla
    @SkanderMoalla
    Jul 14, 2025
    🚀 Big time! We can finally do LLM RL fine-tuning with rewards and leverage offline/off-policy data! ❌ You want rewards, but GRPO only works online? ❌ You want offline, but DPO is limited to preferences? ✅ QRPO can do both! 🧵Here's how we do it:
    Image
    1
  • @SkanderMoalla
    Skander Moalla
    @SkanderMoalla
    Dec 5, 2025
    QRPO is at #NeurIPS2025 in San Diego & #EurIPS in Copenhagen! 🚀 🇺🇸 San Diego: Catch me on Friday, 11 AM – 2 PM, Halls C-E, Poster #512. (I'm on the job market for Fall '26 - RL & LLM post-training! 💼) 🇩🇰 Copenhagen: @matrs01 is still on-site! If you missed his poster on
    Image
  • @SkanderMoalla
    Skander Moalla
    @SkanderMoalla
    Sep 26, 2025
    What's the poster topic to be around LLM RL fine-tuning papers? RL -> Deep RL? RL -> Everything Else? RL -> Batch Offline? (offline/off-policy nature of QRPO)
    @SkanderMoalla
    Skander Moalla
    @SkanderMoalla
    Jul 14, 2025
    🚀 Big time! We can finally do LLM RL fine-tuning with rewards and leverage offline/off-policy data! ❌ You want rewards, but GRPO only works online? ❌ You want offline, but DPO is limited to preferences? ✅ QRPO can do both! 🧵Here's how we do it:
    Image
  • @SkanderMoalla
    Skander Moalla
    @SkanderMoalla
    Sep 19, 2025
    This again shows the limitation of naive online RL in learning rich and robust features. 📉 I was quite surprised to see an Atari agent collapse much faster when given full-resolution observations, due to overfitting its features earlier and then losing plasticity! 😲 We
    @DPyatko
    Daniil Pyatko
    @DPyatko
    Sep 18, 2025
    Did you know that keeping the original colours in Atari makes PPO agents collapse? And, unlike recent research, bigger batch size improves early return but fails at preserving plasticity? We investigate these phenomena and evaluate which interventions are the most effective.
  • @SkanderMoalla
    Skander Moalla
    @SkanderMoalla
    Sep 4, 2025
    A big step for Switzerland 🇨🇭 and a great achievement for our in-house alignment algorithm QRPO (x.com/skandermoalla/…) which has shown remarkable stability and predictability at the 70B scale 🚀!
    @haeggee
    Alex Hägele
    @haeggee
    Sep 2, 2025
    Long in the making, finally released: Apertus-8B and Apertus-70B, trained on 15T tokens of open data from over 1800 languages. Unique opportunity in academia to work on and train LLMs across the full-stack. We managed to pull off a pretraining run with some fun innovations, ...
    Image
    1
Advertisement
Advertisement