Log inSign up
Ce Zhang
47 posts
@cezhhh

Ce Zhang

@cezhhh
CS phd student at UNC Chapel Hill.
Chapel Hill, NC
ceezh.github.io
Joined September 2023
81
Following
99
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @cezhhh
    Ce Zhang
    @cezhhh
    Jul 23
    Excited to share our new work, “Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA.” Grounded LVQA requires a model to answer a question about a long video and localize the short interval supporting its answer. We introduce VideoTreeSearch (VTS), which
    1
  • @cezhhh
    Ce Zhang
    @cezhhh
    Jun 3, 2025
    Recent advances in test-time optimization have led to remarkable reasoning capabilities in LLMs. However, the reasoning capabilities of MLLMs still significantly lag, especially for complex video-language tasks. We present SiLVR, a Simple Language-based Video Reasoning framework.
    Image
    Image
    1
  • @cezhhh
    Ce Zhang
    @cezhhh
    Nov 12, 2024
    Excited to share that LLoVi is accepted to #EMNLP2024. We will present our work in poster session 12, Nov. 14 (Thu.) 14:00-15:30 ET. Happy to have a chat! Check out our paper at: arxiv.org/pdf/2312.17235 Code: github.com/CeeZh/LLoVi Website: sites.google.com/cs.unc.edu/llo…
    @cezhhh
    Ce Zhang
    @cezhhh
    Jan 9, 2024
    Replying to @cezhhh
    First, LLoVi uses a short-term visual captioner to generate textual descriptions of short video clips (0.5-8s in length) densely sampled from a long input video. Afterward, an LLM aggregates the short-term captions to perform long-range temporal reasoning.
    Image
  • @cezhhh
    Ce Zhang
    @cezhhh
    Jan 9, 2024
    Prior long-range video understanding methods are often costly and require specialized long-range video modeling designs. We present LLoVi, a simple yet effective framework with LLM for long-range video question-answering (LVQA). 🧵 arxiv.org/abs/2312.17235 github.com/CeeZh/LLoVi
    Image
    5
Advertisement
Advertisement