π CodeBind (#ACL2026): multimodal alignment that keeps modality-unique info. A shared+specific codebook + compositional VQ aligns 9 modalities, and new modalities plug in the same way.
π visual-ai.github.io/codebind/
π arxiv.org/abs/2605.18257
Computer Vision || Asst. Prof. @HKUniversity | Ex-{Researcher @GoogleAI, Asst. Prof. @BristolUni, Postdoc @Oxford_VGG, @UniofOxford}
Joined July 2018
- π PartCo (#ICML2026): fine-grained category discovery hinges on local parts. It mines part-level correspondence priors from frozen ViT patch tokensβa train-time regularizer that boosts many category discovery methods. π visual-ai.github.io/partco π arxiv.org/abs/2509.22769
- π iVGR (#ICML2026): making MLLMs draw boxes at inference can hurt fine-grained visual reasoning. Dual-stream RL + a consistency reward internalizes grounding into pure textβno boxes at test time, yet SOTA on V*/HRBench. π visual-ai.github.io/ivgr/ π arxiv.org/abs/2605.31096
- 2023: GPT-4 could sketch a partial, static Hong Kong MTR map with some effort. 2026: the full network, live, real-time train positions. πRevisiting transport networks after ~3 years In 2023, we evaluated whether GPT-4 could recreate the Hong Kong MTR, and the result felt remarkable at the time Now, Claude Fable 5 can one-shot a live, interactive MTR map with real-time train positions
- Claude Fable 5 sets a new SOTA on GRAB-lite: 74% @ClaudeDevs are pushing the frontier The best models scored around 20% when GRAB was released On the current trajectory, GRAB-lite is projected to be ~100% by mid-2027 A reminder that tracking the frontier needs hard evals


