1. X
  2. Riashat Islam
Log inSign up
Riashat Islam
615 posts
Image
user avatar
Riashat Islam
@riashatislam
Research Scientist @ms_aifrontiers @MSFTResearch NYC; Ex @HUMAIN @DreamFoldAI PhD @Mila_Quebec, intern @MSRNYC @AppleMLR; RL, Reasoning and LLMs; WorldModels
New York, USA
riashat.github.io
Joined November 2016
1,306
Following
1,827
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Riashat Islam
    @riashatislam
    Apr 25, 2023
    Excited to share that Agent-Controller representations for offline RL in presence of rich exogenous information is now accepted at #ICML2023 (arxiv.org/abs/2211.00164) This is a follow-up of our recent work on latent state discovery (#TMLR'23) arxiv.org/abs/2207.08229
    arxiv.org
    Agent-Controller Representations: Principled Offline RL with Rich...
    Learning to control an agent from data collected offline in a rich pixel-based visual observation space is vital for real-world applications of reinforcement learning (RL). A major challenge in...
  • user avatar
    Riashat Islam
    @riashatislam
    Aug 1
    After Energy-Based Transformers, @AlexiGlad comes back with another really cool work!
    user avatar
    Alexi Gladstone
    @AlexiGlad
    Jul 31
    We discovered a third pretraining axis beyond parameters and data: exploration. Scaling exploration monotonically improves existing models across images/video/language, and unlocks end-to-end generation. In the simplest case, it's just a for loop. Introducing Explorative
    Image
  • user avatar
    Riashat Islam
    @riashatislam
    Jul 4
    Check out our poster on h1 and find @sumeetrm for a chat!
    user avatar
    Sumeet Motwani
    @sumeetrm
    Jul 3
    I'll be presenting 3 papers at ICML 2026🫡 h1 (Spotlight) trains models to reason over longer horizons using curriculum RL over composed short-horizon data. This allows models to generalize to harder tasks and improves performance even at high pass@k. LongCoT isolates and
    Image
  • user avatar
    Riashat Islam
    @riashatislam
    Jun 28
    This is really cool work - you know this work has been thought through when you see @EfroniYonathan on it! Congrats to @Ankur_Samanta_ for the hard work in pulling it off too!
    user avatar
    Ankur Samanta
    @Ankur_Samanta_
    Jun 22
    🚀New work on credit assignment in multi-step reasoning RL post-training🚀 Introducing Self-Reset Policy Optimization (SRPO): i) localize the first wrong reasoning step, ii) reset to that step, iii) learn from counterfactual continuations from there – no external supervision.🧵
    Image
  • user avatar
    Riashat Islam
    @riashatislam
    Jun 6
    A huge loss for the RL, Control and Optimization communities. A legend passed way.
    user avatar
    Subbarao Kambhampati (కంభంపాటి సుబ్బారావు)
    @rao2z
    Jun 5
    Deeply saddened at the passing of my dear colleague, Dimitri Bertsekas. Everyone in RL, OR and control theory already knows of his monumental contributions. Over the past seven years, we at @SCAI_ASU also got to know him as an unwaveringly kind and gracious man of science. He
    Image
    Image

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement