Log inSign up
David McAllister
198 posts
David McAllister profile banner
@davidrmcall

David McAllister

@davidrmcall
PhD Student @berkeley_ai
Joined June 2024
360
Following
1,112
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @davidrmcall
    David McAllister
    @davidrmcall
    Jul 29, 2025
    Excited to share Flow Matching Policy Gradients: expressive RL policies trained from rewards using flow matching. It’s an easy, drop-in replacement for Gaussian PPO on control tasks.
    Image
    00:00
    8
  • @davidrmcall
    David McAllister
    @davidrmcall
    8h
    I’ll present FDFO at ECCV tomorrow during Poster Session 2 (4:30–6:30 PM CEST), in ExHall near poster board #477! 🐈
    @davidrmcall
    David McAllister
    @davidrmcall
    Apr 17
    We developed a simple, sample-efficient online RL technique for post-training image generation models. We see it as a possible steerable alternative to CFG, driven by any scalar reward, including human preference.
    Image
    00:00
    1
  • @davidrmcall
    David McAllister
    @davidrmcall
    Jun 17
    A huge team effort! Fully open data, infrastructure and eval for the community to build on
    Image
    00:00
    @ritvik_singh9
    Ritvik Singh
    @ritvik_singh9
    Jun 17
    Image
    02:51
    Introducing ABC: open data, training, and infrastructure for robotics. We release the largest teleop dataset to date, and extensively investigate design decisions, pretraining, and post-training techniques. @arthurallshire @Cinnabar233 @adamrasb @redstone_hong @davidrmcall
    5
  • @davidrmcall
    David McAllister
    @davidrmcall
    May 25
    Robot (@arthurallshire)
    Image
    00:00
    2
  • @davidrmcall
    David McAllister
    @davidrmcall
    Apr 17
    We developed a simple, sample-efficient online RL technique for post-training image generation models. We see it as a possible steerable alternative to CFG, driven by any scalar reward, including human preference.
    Image
    00:00
    11
Advertisement
Advertisement