Excited to share our new work, “Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA.”
Grounded LVQA requires a model to answer a question about a long video and localize the short interval supporting its answer. We introduce VideoTreeSearch (VTS), which
CS phd student at UNC Chapel Hill.

