There has been a clear trend in the last months moving from VLA-type approaches to Video Generative Models + Inverse Dynamics Models (VAM).
While the probable main reason of this recent growth is the latest improvements in video generative models, I believe this shift is
Robotics Tinkerer. RS @Amazon FAR
Prev: @META (FAIR), @DFKI, @TUDarmstadt
robotgradient.com X github.com/robotgradient
Joined November 2017
- While I really liked the article, it feels to me that this physical commonsense can be better capture by predicting next observations (i.e. world models) and planning on it, rather than training a policy on predicting next action (i.e. behavioral cloning)
- This was very challenging and very cool to see evolve! I personally was no sure if it would work, but @irmakkguzey pushed so hard to show it does. Learning dexterous robot policies with only human video data, using the egocentric view from Aria2 glasses, chill and easy 😁Dexterous manipulation by directly observing humans - a dream in AI for decades - is hard due to visual and embodiment gaps. With simple yet powerful hardware - Aria 2 glasses 👓 - and our new work AINA 🪞, we are now one significant step closer to achieving this dream.
- The expert mode is going to bring a lot of news in the future 🙃NEO The Home Robot Order Today
- Anyone interested in tactile sensing for robotics should be following Akash's solid releases. How should we integrate rich tactile sensing modality for policy learning?Robots need touch for human-like hands to reach the goal of general manipulation. However, approaches today don’t use tactile sensing or use specific architectures per tactile task. Can 1 model improve many tactile tasks? 🌟Introducing Sparsh-skin: tinyurl.com/y935wz5c 1/6






