🎇Thrilled to release VLM^3! Most 3D vision papers nowadays still spend months/years designing complex archs/losses/augmentations for different tasks. Are they necessary? VLM^3 shows that most designs that you think are important for 3D vision are [not] important at all!
We just released MobileMoE, first sub-B-active-parameter MoE language model family.
MoE isn't just for 100B+ parameter models on servers. At sub-B scale, sparse expert routing lets you match dense models at 2-4x fewer FLOPs while fitting in mobile DRAM.
arxiv.org/pdf/2605.27358
Grateful to @sallywf and @EETimes for the thoughtful writeup of my Embedded Vision Summit keynote. The thesis in one line: the next decade of AI won't be won by the biggest model, but by the smartest, most efficient one that lives on the devices you wear!
Audio is the most ignored perception modality in on-device AI.
Every smart glass, robot, and drone has a mic. Almost none fuse audio + vision at perception time.
Vision-only is the vibe-coded version of multimodal perception.