Log inSign up
Yunze Man
125 posts
Yunze Man profile banner
@yunzeman

Yunze Man

@yunzeman
Research Scientist, NVIDIA GEAR. Training robotics foundation model. Ph.D in UIUC
Santa Clara, CA
yunzeman.github.io
Joined October 2019
275
Following
433
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @yunzeman
    Yunze Man
    @yunzeman
    Aug 25
    It is fun to see so many companies committed to in-context learning / video prompting and making progress there. Video naturally carries rich and nuanced supervision that a sentence. So video and language conditioning should go hand in hand rather than compete. They both matter
    @SkildAI
    Skild AI
    @SkildAI
    Aug 25
    Introducing S1, our new foundation model that learns from one example. It can be taught 10-minute long tasks that it has never seen before, from one video prompt without any fine-tuning. Watch S1 operate in real-time via in-context learning:
    Image
    00:00
  • @yunzeman
    Yunze Man
    @yunzeman
    Aug 9
    Q and V are very policy-dependent. Pre-training them feels not applicable, at least at the moment. The more scalable approach is to pre-train a rewarder, dense plus sparse. This was true for LLM, also probably true for robotics?
    @chelseabfinn
    Chelsea Finn
    @chelseabfinn
    Aug 9
    Pretraining a Q-function often doesn’t actually help RL finetuning, compared to initializing Q from scratch. We find that pretraining Q-functions on data from diverse policies is critical to see improvements from pretraining. Paper: arxiv.org/abs/2607.27203
  • @yunzeman
    Yunze Man
    @yunzeman
    Aug 5
    To have an arena called “Vision Arena” is crazy. Are we talking about video/image understanding? tracking? segmentation? Claude the best? anyone working with video knows Claude is way worse than Gemini (G is also very bad, btw) It would make more sense if you break it down.
    @arena
    Arena.ai
    @arena
    Aug 3
    Replying to @arena
    Qwen3.8-Max ranks #2 in Vision Arena scoring 1,305. Second only to Claude Fable 5 (High) which has only a 13pt lead.
    Image
  • @yunzeman
    Yunze Man
    @yunzeman
    Jul 27
    I resonate with this post deeply. Agents keep getting faster, and I increasingly feel myself being the productivity bottleneck: all the high-level decisions, managing so many parallel agents, understanding the output Health matters more in AI research. Wishing Lilian the best.
    @lilianweng
    Lilian Weng
    @lilianweng
    Jul 27
    It is a hard and sad decision. I shared this message with folks at Thinky. Thank you all for the time together♥️ Just as the last sentence in my message: The future worth building is human.
    Image
  • @yunzeman
    Yunze Man
    @yunzeman
    Jul 23
    Robot dat scaling usually means more tasks, scenes, and hours. GEN-1 is scaling over EEFs too. I love how data-first they are: use all feedback and signal to see what type of data is missing, collect it, and iterate. That flywheel is exactly what we need to solve robotics.
    @GeneralistAI
    Generalist
    @GeneralistAI
    Jul 23
    Why create robot intelligence for just one hand, when we could have it learn from many? GEN-1, our latest embodied foundation model, now supports a broad range of end effectors from 5-finger hands, to specialized tools, and everything in between.
    Image
    00:00
Advertisement
Advertisement