Log inSign up
Vikas Chandra
348 posts
@vikasc

Vikas Chandra

@vikasc
Senior Director of #AI Research @Meta | CMU Ph.D. | Ex visiting faculty at Stanford
Menlo Park, CA
v-chandra.github.io
Joined April 2009
204
Following
629
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @vikasc
    Vikas Chandra
    @vikasc
    Aug 19
    The video of my keynote talk at this year's @EmbVisionSummit is now published. Scaling Down is the New Scaling Up youtu.be/amdqNKnNL-s
    1
  • @vikasc
    Vikas Chandra
    @vikasc
    Jun 1
    Vision Language Models are native 3D learners!
    @cai_zhipeng
    Zhipeng Cai
    @cai_zhipeng
    Jun 1
    🎇Thrilled to release VLM^3! Most 3D vision papers nowadays still spend months/years designing complex archs/losses/augmentations for different tasks. Are they necessary? VLM^3 shows that most designs that you think are important for 3D vision are [not] important at all!
    Image
    1
  • @vikasc
    Vikas Chandra
    @vikasc
    May 29
    We just released MobileMoE, first sub-B-active-parameter MoE language model family. MoE isn't just for 100B+ parameter models on servers. At sub-B scale, sparse expert routing lets you match dense models at 2-4x fewer FLOPs while fitting in mobile DRAM. arxiv.org/pdf/2605.27358
    Image
  • @vikasc
    Vikas Chandra
    @vikasc
    May 20
    Grateful to @sallywf and @EETimes for the thoughtful writeup of my Embedded Vision Summit keynote. The thesis in one line: the next decade of AI won't be won by the biggest model, but by the smartest, most efficient one that lives on the devices you wear!
    Image
    Meta: Scaling Agentic AI on Resource-Constrained Devices- EE Times
    From eetimes.com
    1
  • @vikasc
    Vikas Chandra
    @vikasc
    May 2
    Audio is the most ignored perception modality in on-device AI. Every smart glass, robot, and drone has a mic. Almost none fuse audio + vision at perception time. Vision-only is the vibe-coded version of multimodal perception.
Advertisement
Advertisement