Introducing "Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection"
🤖 VLMs answer well—but only if the evidence is already visible. How can we bring this capability to real agent scenarios, where the current view is insufficient to answer directly?