It is fun to see so many companies committed to in-context learning / video prompting and making progress there.
Video naturally carries rich and nuanced supervision that a sentence. So video and language conditioning should go hand in hand rather than compete. They both matter
Research Scientist, NVIDIA GEAR. Training robotics foundation model. Ph.D in UIUC






