Log inSign up
Baifeng
427 posts
Baifeng profile banner
@baifeng_shi

Baifeng

@baifeng_shi
@physical_int, @berkeley_ai, vision, robotics; ex @nvidia; opinions are my own.
bfshi.github.io
Joined September 2021
696
Following
2,338
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @baifeng_shi
    Baifeng
    @baifeng_shi
    Mar 24
    Humans can see in high-res, high-FPS in real-time. Why can't VLMs? Introducing AutoGaze: ViTs/VLMs "gaze" only at key video regions! Up to 4-100x token savings, 19x speedup, and enables scaling to 4K-res 1K-frame videos. 📄 arxiv.org/abs/2603.12254 🌐 autogaze.github.io 🤗
    Image
    00:00
    47
  • @baifeng_shi
    Baifeng
    @baifeng_shi
    Jun 20
    Learning from task-agnostic, explorative experience!
    @junyi42
    Junyi Zhang
    @junyi42
    Jun 19
    Children learn from play. Can robots do the same? We propose 𝐏𝐥𝐚𝐲𝐟𝐮𝐥 𝐀𝐠𝐞𝐧𝐭𝐢𝐜 𝐑𝐨𝐛𝐨𝐭 𝐋𝐞𝐚𝐫𝐧𝐢𝐧𝐠, a paradigm that gives embodied coding agents a play stage before downstream tasks arrive, and instantiate it with 𝐑𝐀𝐓𝐬 (Robotics Agent Teams), where robots
    Image
    00:00
  • @baifeng_shi
    Baifeng
    @baifeng_shi
    Apr 29
    Congrats on the release!
    @leoyerrrr
    Hanrong YE @ Nvidia
    @leoyerrrr
    Apr 28
    @nvidia introduces Nemotron 3 Nano Omni LLMs. A year ago, we started research on omni-modal LLMs—covering everything from architecture to data—and released OmniVinci (ICLR 2026). We’ve received some feedback since then, but two clear themes emerged: 9B parameters weren't quite
    Image
  • @baifeng_shi
    Baifeng
    @baifeng_shi
    Apr 17
    Our newest generalist model with great zero-shot capabilities! Some highlights: - can utilize both high-quality and low-quality data via proper conditioning - careful design of diverse data - can also condition on imagined goal image from an image world model - as a result, the
    @physical_int
    Physical Intelligence
    @physical_int
    Apr 16
    Our newest model, π0.7, has some interesting emergent capabilities: it can control a new robot to fold shirts for which we had no shirt folding data, figure out how to use an appliance with language-based coaching, and perform a wide range of dexterous tasks all in one model!
    Image
    00:00
    2
  • @baifeng_shi
    Baifeng
    @baifeng_shi
    Apr 9
    Check out @LongTonyLian’s cool parallel reasoning work!
    @LongTonyLian
    Long Lian
    @LongTonyLian
    Apr 8
    Our parallel reasoning project ThreadWeaver is now open-sourced 🎉! Check out our Data Gen/SFT/RL recipe at github.com/facebookresear… In case you don't know, ThreadWeaver 🧵⚡️ is the first parallel reasoning method to achieve comparable reasoning performance to widely-used
Advertisement
Advertisement