Log inSign up
Mingfei Chen
56 posts
@lasiafly

Mingfei Chen

@lasiafly
PhD Student @UW | PhD Fellow @Google | Multi-Modal LLMs, Spatial AI, XR, Robotics | Previously @Meta & @NUSingapore
Seattle, WA
mingfeichen.com
Joined December 2021
201
Following
114
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @lasiafly
    Mingfei Chen
    @lasiafly
    Jul 22
    We may not need a multibillion-parameter LLM to understand what we see and hear—and where it is in 3D. Introducing SceneBind, the first omni-modal representation to jointly model semantics and explicit 3D spatial information across vision, binaural audio, and language in
    Image
    00:00
    3
  • @lasiafly
    Mingfei Chen
    @lasiafly
    Jun 18
    Glad that EgoMAN has been provisionally accepted at #ECCV2026. EgoMAN Connects vision–language reasoning with physically grounded 6-DoF hand motion for stage-aware trajectory prediction from egocentric videos—enabling proactive AI interaction, human motion synthesis, and robotic
    Image
    00:00
    6
  • @lasiafly
    Mingfei Chen
    @lasiafly
    Jun 18
    Spatial-temporal memory will be more and more important if agentic robotic system or long-term action modeling is trending. We need to know what and where and be able to plan over KEY memory
    @lucacarlone1
    Luca Carlone
    @lucacarlone1
    Jun 17
    great to see our work "DAAAM: Describe Anything Anywhere At Any Moment" with Nicolas Gorlo and Lukas Schmid featured on the MIT News: news.mit.edu/2026/could-ai-… if you are interested in robot memory, scene graphs, mapping, or question answering, check it out! #mitSparkLab #CVPR2026
  • @lasiafly
    Mingfei Chen
    @lasiafly
    Jun 17
    Cool language guided point tracking work. I’m increasingly convinced motion tracks may be a better state for robotic world/action models than raw video/pixel tokens. But do we need a large MLLM, or can a lightweight LLM that parses text query -> 3D-tracker interface be enough,
    @allen_ai
    Ai2
    @allen_ai
    Jun 17
    We're releasing MolmoMotion, a 3D motion forecasting model. Given one or a few video frames, 3D points on an object, & an instruction like "Put the white bowl on the table," MolmoMotion predicts where those points will go over the next few seconds in a shared 3D world frame. 🧵
    Image
    00:00
  • @lasiafly
    Mingfei Chen
    @lasiafly
    Jun 17
    Qwen-RobotNav (qwen.ai/blog?id=qwen-r…) is an exciting step toward real embodied agentic systems. What I like most is the deployable design: upper-level LLM planner + multimodal navigation agent + lightweight action head Instead of forcing one giant model to do everything at
    Image
Advertisement
Advertisement