I’m returning to Berkeley as a Ph.D. in EECS, specializing in Computer Vision!
I am grateful for my mentors/family/friends who have guided me through this process. I cant wait to continue building exciting work in @berkeley_ai. Stay tuned!!
We are presenting "Generate, but Verify". Our backtracking schema reduces VLM hallucinations significantly.
See how we implemented this in a single VLM architecture.
Visit our poster session to chat with us!
📄: arxiv.org/abs/2504.13169@NeurIPSConf
Excited to see @deepseek_ai push the generation-verification idea to IMO-level math (insane!)
We’ve been exploring the same flow earlier this year and it really cuts VLM hallucinations!
NeurIPS folks, come chat at our poster: Thu Dec 4, 4:30–7:30 pm, Exhibit Hall C/D/E, #4618.
Me as an undergrad, I am super fortunate to learn and work w Patrick - has so many ideas in bringing up exciting problems in AI that needs to be solved someday
and this essay is amazing - check it out!
Accidentally wrote a blog post on dynamic human-AI interaction this week, sharing some tentative ideas I find interesting in this field and the connection to our REVERSE-VLM. I’ll be presenting it at NeurIPS next week.
Happy to chat more at San Diego 🙂
patrickthwu.com/posts/dynamic-…
Attending EMNLP2025 (SuZhou)!
🍀 Showing our work visual puzzles with @_dmchan , and supporting the community as a volunteer
Happy to meet everyone and to make new connections!
@emnlpmeeting#EMNLP2025arxiv.org/abs/2505.23759 : An exciting dataset to test vlm capabilities
🔥 Excited to share that my first paper “Puzzled by Puzzles: When Vision-Language Models Can’t Take a Hint” has been accepted to main conference #EMNLP2025 🎉
📝 Paper: arxiv.org/abs/2505.23759
Follow the thread below and stay tuned if interested