We may not need a multibillion-parameter LLM to understand what we see and hear—and where it is in 3D.
Introducing SceneBind, the first omni-modal representation to jointly model semantics and explicit 3D spatial information across vision, binaural audio, and language in
PhD Student @UW | PhD Fellow @Google | Multi-Modal LLMs, Spatial AI, XR, Robotics | Previously @Meta & @NUSingapore



