Log inSign up
Arijit Ray
82 posts
Arijit Ray profile banner
@ARRay693

Arijit Ray

@ARRay693
Research Scientist @Google | Teaching AI to complete tasks that humans need but don’t want to do.
SF, CA
arijitray.com
Joined November 2015
814
Following
193
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @ARRay693
    Arijit Ray
    @ARRay693
    Feb 18
    "It is by logic that we prove, but by [abstract] intuition that we discover." - Henri Poincaré. When faced with a complex problem, we pause, we think. Not exactly in words, not exactly in images — in something more abstract, something harder to name. So, for truly
    Image
    00:00
    2
  • @ARRay693
    Arijit Ray
    @ARRay693
    Nov 10, 2025
    Game your benchmark first (and de-bias it) before others do!
    @_ellisbrown
    Ellis Brown
    @_ellisbrown
    Nov 10, 2025
    🌶️ hot take 🌶️ > we should normalize training on the test set yes, you read that right. no, I'm not joking. and, yes... I have taken ML 101 👉 here's why this is crucial for future multimodal LLM research [1/n] 🧵
  • @ARRay693
    Arijit Ray
    @ARRay693
    Nov 7, 2025
    SIMS-V offers free (simulated) rich accurate video annotations for object relationships, distances, and temporal tracking—capabilities often lacking in existing video training datasets. 🎞️💫 Mix it into your data and boost your model's performance on video reasoning tasks! Code
    @_ellisbrown
    Ellis Brown
    @_ellisbrown
    Nov 7, 2025
    MLLMs are great at understanding videos, but struggle with spatial reasoning—like estimating distances or tracking objects across time. the bottleneck? getting precise 3D spatial annotations on real videos is expensive and error-prone. introducing SIMS-V 🤖 [1/n]
    Image
    00:00
  • @ARRay693
    Arijit Ray
    @ARRay693
    Nov 7, 2025
    We live, feel, and create by perceiving the world as visual spaces unfolding through time — videos. Our memories and even our language are spatial: mind-palaces, mind-maps, "taking steps in the right direction..." Super excited to see Cambrian-S pushing this frontier! And,
    @sainingxie
    Saining Xie
    AMI Labs
    @sainingxie
    Nov 7, 2025
    Introducing Cambrian-S it’s a position, a dataset, a benchmark, and a model but above all, it represents our first steps toward exploring spatial supersensing in video. 🧶
    Image
    00:00
    1
  • @ARRay693
    Arijit Ray
    @ARRay693
    Oct 5, 2025
    Indeed! Come Tuesday morning to our poster. Super excited to chat about multi-step vision-language reasoning and how SAT, and simulations/world models can teach this to Multimodal Language models.
    @DJiafei
    Jiafei Duan
    @DJiafei
    Oct 5, 2025
    SAT will be presented at #COLM2025 by @ARRay693 ! Go talk to him and learn more about SAT!
Advertisement
Advertisement