1. X
  2. Arijit Ray
Log inSign up
Arijit Ray
82 posts
Image
user avatar
Arijit Ray
@ARRay693
AI PhD Student, w/ Profs @kate_saenko_, @RanjayKrishna | Teaching machines to reason better in the digital and physical world | Prev @Google, @AIatMeta
Cambridge, MA
arijitray.com
Joined November 2015
804
Following
192
Followers
RepliesRepliesMediaMedia
  • Pinned
    user avatar
    Arijit Ray
    @ARRay693
    Feb 18
    "It is by logic that we prove, but by [abstract] intuition that we discover." - Henri Poincaré. When faced with a complex problem, we pause, we think. Not exactly in words, not exactly in images — in something more abstract, something harder to name. So, for truly
    Image
    00:00
  • user avatar
    Arijit Ray
    @ARRay693
    Nov 10, 2025
    Game your benchmark first (and de-bias it) before others do!
    user avatar
    Ellis Brown
    @_ellisbrown
    Nov 10, 2025
    🌶️ hot take 🌶️ > we should normalize training on the test set yes, you read that right. no, I'm not joking. and, yes... I have taken ML 101 👉 here's why this is crucial for future multimodal LLM research [1/n] 🧵
  • user avatar
    Arijit Ray
    @ARRay693
    Nov 7, 2025
    SIMS-V offers free (simulated) rich accurate video annotations for object relationships, distances, and temporal tracking—capabilities often lacking in existing video training datasets. 🎞️💫 Mix it into your data and boost your model's performance on video reasoning tasks! Code
    user avatar
    Ellis Brown
    @_ellisbrown
    Nov 7, 2025
    MLLMs are great at understanding videos, but struggle with spatial reasoning—like estimating distances or tracking objects across time. the bottleneck? getting precise 3D spatial annotations on real videos is expensive and error-prone. introducing SIMS-V 🤖 [1/n]
    Image
    00:00
  • user avatar
    Arijit Ray
    @ARRay693
    Nov 7, 2025
    We live, feel, and create by perceiving the world as visual spaces unfolding through time — videos. Our memories and even our language are spatial: mind-palaces, mind-maps, "taking steps in the right direction..." Super excited to see Cambrian-S pushing this frontier! And,
    user avatar
    Saining Xie
    AMI Labs
    @sainingxie
    Nov 7, 2025
    Introducing Cambrian-S it’s a position, a dataset, a benchmark, and a model but above all, it represents our first steps toward exploring spatial supersensing in video. 🧶
    Image
    00:00
  • user avatar
    Arijit Ray
    @ARRay693
    Oct 5, 2025
    Indeed! Come Tuesday morning to our poster. Super excited to chat about multi-step vision-language reasoning and how SAT, and simulations/world models can teach this to Multimodal Language models.
    user avatar
    Jiafei Duan
    @DJiafei
    Oct 5, 2025
    SAT will be presented at #COLM2025 by @ARRay693 ! Go talk to him and learn more about SAT!

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
Advertisement
Advertisement