Full of childlike wonder. Teaching robots manners. RL Lead @ Apptronik. UT Austin PhD candidate. Past: Boston Dynamics AI Institute, NASA JPL, MIT โ20.
love this-
RL with dense guidance isnโt really RL at all. weโre still pretty bad at efficient exploration, but works like this, RFCL (@Stone_Tao), OmniReset (@patrickhyin, @ty_westenbroek), and others are showing great progress forwards!
๐๐ผ๐ผ๐ฑ ๐บ๐ฎ๐ป๐ถ๐ฝ๐๐น๐ฎ๐๐ถ๐ผ๐ป ๐ฝ๐ผ๐น๐ถ๐ฐ๐ถ๐ฒ๐ ๐บ๐ฎ๐ ๐๐๐ฎ๐ฟ๐ ๐๐ถ๐๐ต ๐ฏ๐ฒ๐๐๐ฒ๐ฟ ๐๐๐ฎ๐๐ฒ๐, ๐ป๐ผ๐ ๐ฏ๐ฒ๐๐๐ฒ๐ฟ ๐ฟ๐ฒ๐๐ฎ๐ฟ๐ฑ๐.
We sample diverse, physically feasible contact states as starts + goals for RL.
The behaviors that emerge are surprisingly dynamic and
this is โtraining an RL policy for inverse kinematicsโ all over again. itโs just tau = -g(q).
the trade off is data vs modeling the mass and inertias. I would argue modeling mass is easier (and inherently more robust), but still always cool seeing stuff work on hardware.
Havenโt compared to test-time RTC, itโs a bit hard to both implement and debug๐ฅฒ
The policy here is actually trained on 30hz but i just bumped up the execution hz to 50 without issue. Since this is just 1 robot, inference just runs as fast as it can (on Armory!) and the prefix