Log inSign up
David Fan
AMI Labs
191 posts
@DavidJFan

David Fan

AMI Labs
@DavidJFan
AMI Labs | ex-Meta FAIR | @Princeton CS '19 Building the next revolution of AI models that understand the real world.
New York City
scholar.google.com/citations?user…
Joined June 2013
509
Following
1,799
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @DavidJFan
    David Fan
    AMI Labs
    @DavidJFan
    Mar 4
    [1/9] What happens when you treat vision as a first-class citizen during multimodal pretraining? To find out, we studied the design space of training Transfusion-style models that input and output all modalities, from scratch. Here is what we learned about visual representations,
    arXiv logo
    arxiv.org
    Beyond Language Modeling: An Exploration of Multimodal Pretraining
    The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models remains opaque....
    12
  • @DavidJFan
    David Fan
    AMI Labs
    @DavidJFan
    Aug 7
    Nice work Junlin!!
    @han_junlin
    Junlin Han
    @han_junlin
    Aug 6
    We live in a multimodal world. We see, talk, act, and dream. Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch? Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and
    Image
    1
  • @DavidJFan
    David Fan
    AMI Labs
    @DavidJFan
    Jul 26
    Nice work!! Really glad to hear WebSSL works well for robot policies 😁 Representation matters
    @jeffacce
    Jeff Cui
    @jeffacce
    Jul 22
    Your policy doesn't need 7B params. It simply needs dense features. Introducing Patch Policy: pretrained ViT + small transformer beats OpenVLA-OFT with 0.7% of its params, and trains on a 5090. Here it inserts a cable (~2mm tol), and does it again as we unplug mid-rollout. 🧵
    Image
    00:00
    2
  • @DavidJFan
    David Fan
    AMI Labs
    @DavidJFan
    Jul 7
    @__JohnNguyen__ and I are presenting the Beyond Language Modeling paper as an ICML spotlight in 30 minutes at the 10:30 AM poster session! I’ll also be at the AMI Mixer on Thursday and hanging around in Seoul until early next week :)
    @DavidJFan
    David Fan
    AMI Labs
    @DavidJFan
    Mar 4
    [1/9] What happens when you treat vision as a first-class citizen during multimodal pretraining? To find out, we studied the design space of training Transfusion-style models that input and output all modalities, from scratch. Here is what we learned about visual representations,
  • @DavidJFan
    David Fan
    AMI Labs
    @DavidJFan
    Apr 16
    Congrats Dr. @TongPetersb!! Your research journey has clearly culminated in a very cohesive and inspiring narrative that unifies several areas of work with much scope for future expansion :D It's a testament to your work ethic, good taste in problems, and attention to detail. I'm
    @__JohnNguyen__
    John Nguyen
    AMI Labs
    @__JohnNguyen__
    Apr 16
    Congrats Dr. Tong! Really glad to be a part of your PhD journey @TongPetersb
    Image
    Image
    Image
    00:00
Advertisement
Advertisement