Action chunking is a critical component in virtually all modern approaches to imitation learning for robotics.
But why is it so critical, and do we really need action chunking? Check out our latest work to find out! (1/n)
action-chunking.github.io
MTS @physical_int | Incoming Assistant Professor @UTAustin CS | Previously: Postdoc @Berkeley_EECS, PhD @uwcse
Joined May 2025
- How can we leverage VLAs to learn to solve complex tasks that fall outside their typical capabilities? We introduce Semantic Action RL—treat the VLA’s language prompt as an action and optimize this with RL! Check out Jagdeep's thread for all the details! semantic-action-rl.github.ioHow can generalist policies adapt to new challenges at deployment using skills they already have? We optimize VLA *prompt inputs* with reinforcement learning, enabling efficient real-robot adaptation on complex tasks where existing methods struggle. 🧵 semantic-action-rl.github.io
- How can we elicit useful, semantically meaningful behaviors from generalist policies? We introduce Flow Reversal Steering (FRS) as a method to refine coarse, semantically meaningful commands into effective robot actions! flow-reversal-steering.github.io 1/N
- Come check out our workshop on post-training robot foundation models at RSS 2026! Also consider participating in our real-world RL challenge!#RSS2026 Call for participants 📢 Excited to announce our RSS 2026 Workshop: Post-Training for Robotics Foundation Models, together with the first Real-World Reinforcement Learning Challenge! The workshop is held on July 13 in Sydney. posttraining-for-robotics.github.io。
- Reliable rewards are critical for effective RL, yet in most robotic applications obtaining such rewards requires significant task-specific human effort. Can we do better? Check out RoboReward, our new generalist, language-conditioned reward model for real-world robot RL!Reliable rewards are a bottleneck for real-world RL for robotics: human labels are costly, and handcrafted rewards are brittle. In RoboReward 🤖💰, we study VLMs as reward models and find they are unreliable across tasks, embodiments, and scenes. Paper: arxiv.org/abs/2601.00675




