Portrait of Jianshu Zhang

Jianshu Zhang 张鉴殊

CS Ph.D. Student Northwestern University andyzhang [at] u.northwestern.edu

About Me

Hi! I’m Jianshu Zhang (张鉴殊), a second-year CS Ph.D. student at Northwestern University. My research focuses on building multimodal and agentic AI systems that are truly ready to understand, reason about, and act in the interactive and physical world. Feel free to reach out if you’d like to chat.

Research Interests

How can large vision-language models truly step into our world?

  • Safe and diverse behavior: before VLMs act in open environments, they need to remain safe, reliable, and robust across diverse users, contexts, and failure modes.
  • Clear visual perception: VLMs must see the world precisely, preserving fine-grained perception and visual grounding instead of relying mainly on language priors.
  • Spatial and physical intelligence: beyond recognizing images, VLMs should build spatial mental models, reason about unseen or interactive environments, and connect perception to physical action.

News

  • [Sep 2026] ProgressCompass was released, showing that embodied progress reward models are not blind but lost without the right context, and supplying that context with an agentic loop.
  • [Sep 2026] Video2Skill was released, asking whether VLMs can turn streaming visual experience into a persistent library of reusable embodied skills.
  • [Jul 2026] Progress Reward Modeling for Robotic Learning was released, providing a comprehensive survey of progress reward modeling for robotic learning.
  • [Jul 2026] I received the Outstanding Reviewer Award at ACL 2026.
  • [May 2026] SpaceNum was released, revisiting spatial-numerical understanding in VLMs.
  • [May 2026] ProgressLM was selected as Oral in ACL 2026.
  • [May 2026] I was selected as a Silver Reviewer in ICML 2026.
  • [Apr 2026] AdvEvo-MARL was accepted to ICML 2026.
  • [Apr 2026] ProgressLM and COREWeaver were accepted to ACL 2026 (Main).
Show more
  • [Jan 2026] MindCube was accepted to ICLR 2026.
  • [Jan 2026] ProgressLM was released, probing how well VLMs reason in progress.
  • [Sep 2025] Honored to receive McCormick School of Engineering Fellowship from Northwestern University.
  • [Sep 2025] Joined Northwestern University as a Ph.D. student in Computer Science.
  • [Aug 2025] Evo-MARL and FairReason were accepted to T2FM@ICCV 2025.
  • [Aug 2025] WebCoT was accepted to EMNLP 2025 (Findings) and AIA@COLM 2025.
  • [Jun 2025] MultiVerse was accepted to ICCV 2025.
  • [Jun 2025] Graduated with a B.E. degree and selected as Outstanding Undergraduate.
  • [May 2025] VLM2-Bench was accepted to ACL 2025 (Main), and Bridge-Coder to ACL 2025 (Findings).
  • [May 2025] Honored with the Lei Jun Computer Breakthrough Award (50K RMB).
  • [May 2025] CAN was accepted to ICML 2025.
  • [Jan 2025] PVIT was accepted to ICLR 2025.
  • [Nov 2024] PVIT-3M dataset ranked Top 3 in downloads on Hugging Face.
  • [Oct 2024] Awarded the National Scholarship (top 0.2% nationally).
  • [Sep 2024] Image Textualization accepted to NeurIPS 2024 (D&B).
  • [Sep 2024] MLLM-Protector and FIRST accepted to EMNLP 2024 (Main).
  • [Mar 2024] CORE accepted to CogSci 2024 (Oral).
  • [Dec 2023] FuzzLLM accepted to ICASSP 2024.

Publications

2026

  1. ProgressCompass: a robot episode plays with the step instruction supplied as context, while the frozen PRM drifts and ProgressCompass follows the true progress New
    Jianshu Zhang*, Keliang Wu*, Chengxuan Qian, Xiyuan Yang, Ce Zhang, Ariel Tian, Anbang Liu, Haoran Lu, Han Liu
    arXiv preprint arXiv:2609.36684
  2. Video2Skill: videos stream in, events reuse known skills or create new ones in a persistent skill library New
    Jianshu Zhang*, Ce Zhang*, Xiyuan Yang, Chenwei Xu, Haoran Lu, Yijiang Li, Yaqi Xie, Katia P. Sycara, Han Liu
    arXiv preprint arXiv:2609.36691
  3. Progress Reward Modeling survey: a sparse success signal only fires at the end while a progress reward gives feedback every step; progress models can output a score, a delta, a ranking, or reward code New
    Jianshu Zhang*, Keliang Wu*, Haoran Lu*, Anbang Liu, Ce Zhang, Weijie Yin, Chengxuan Qian, Xiyuan Yang, Zhenyu Pan, Guo Ye, Han Liu
    arXiv preprint arXiv:2607.21655
  4. SpaceNum: a VLM guesses rotation angles and object coordinates like a roulette wheel; across 18 VLMs the best averages 39.8% vs 30.0% random guess New
    Jianshu Zhang*, Yijiang Li*, Huifeixin Chen, Haoran Lu, Letian Xue, Bingyang Wang, Han Liu
    New Released
  5. ProgressLM: a VLM retrieves the matching demo frame (40%), then mentally simulates the remaining change to estimate 46% progress, instead of a direct guess of about 70% ACL
    Jianshu Zhang*, Chengxuan Qian*, Haosen Sun, Haoran Lu, Dingcheng Wang, Letian Xue, Han Liu
    ACL 2026 (Main, Oral); ICLR 2026 Workshop on World Models
  6. MindCube: four limited views fill a cognitive map, the model reasons over it to find unseen objects, and map-then-reason training lifts accuracy from 37.8% to 61.3% ICLR
    Baiqiao Yin, Qineng Wang, Pingyue Zhang, Jianshu Zhang, Kangrui Wang, Zihan Wang, Jieyu Zhang, Keshigeyan Chandrasegaran, Han Liu, Ranjay Krishna, Saining Xie, Manling Li, Jiajun Wu, Li Fei-Fei
    ICLR 2026; SP4V@ICCV 2025 (Best Paper Award)
  7. Explore to Evolve: an agent explores web pages, files and charts for evidence, then evolves an aggregation program into a verifiable research question ACL
    Rui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, Dong Yu, Kam-Fai Wong
    ACL 2026 (Main)
  8. AdvEvo-MARL: attacker and defender agents co-evolve over rounds until defenders block evolving jailbreak attempts, cutting worst-case attack success from 38% (baselines) to 18% ICML
    Zhenyu Pan, Yiting Zhang, Zhuo Liu, Yolo Yunlong Tang, Zeliang Zhang, Haozheng Luo, Yuwei Han, Jianshu Zhang, Dennis Wu, Hong-Yu Chen, Haoran Lu, Haoyang Fang, Manling Li, Chenliang Xu, Philip S. Yu, Han Liu
    ICML 2026
  9. MagicSim: language commands flow through Command, Skill, Planner, Robot and Record as a humanoid grasps, pours and hands over a cup, then one batched runtime serves RL benchmarks, data collection and VLM agents arXiv
    Haoran Lu, Songling Liu, Yue Chen, Guo Ye, Mutian Shen, Shuyang Yu, Yu Xiao, Jihai Zhao, Shang Wu, Jianshu Zhang, Xiangtian Gui, Chuye Hong, Yuran Wang, Maojiang Su, Jiayi Wang, Ruihai Wu, Zhaoran Wang, Han Liu
    arXiv preprint arXiv:2606.17511
  10. Phys4D: for the same prompt, base Wan2.2 makes a rolling ball grow and split while Phys4D keeps one ball on a smooth path, with consistent depth, motion and 4D trajectories; Physics-IQ rises from 21.3 to 35.8 arXiv
    Haoran Lu, Shang Wu, Songling Liu, Jianshu Zhang, Maojiang Su, Guo Ye, Chenwei Xu, Lie Lu, Pranav Maneriker, Fan Du, Manling Li, Zhaoran Wang, Han Liu
    arXiv preprint arXiv:2603.03485

2025

  1. VLM2-Bench: linking matching visual cues across images (same shoes, same plush toy) is easy for people (94.4%) but hard for VLMs (best 59.6%) ACL
    Jianshu Zhang*, Dongyu Yao*, Renjie Pi, Paul Pu Liang, Yiren Fung
    ACL 2025 (Main)
  2. MultiVerse: a VLM is graded with per-turn checklist items across a multi-turn conversation about an image, and even the best model (GPT-4o) scores under 50% ICCV
    Young-Jun Lee, Byung-Kwan Lee, Jianshu Zhang, Yechan Hwang, Byungsoo Ko, Han-Gyu Kim, Dongyu Yao, Xuankun Rong, Eojin Joo, Seung-Ho Han, Bowon Ko, Ho-Jin Choi
    ICCV 2025; ICCV 2025 Workshop KnowledgeMR
  3. PVIT: given a personalized prefix of faces and names, the model finds Kate in a new scene and refuses to answer about people it was never introduced to ICLR
    Renjie Pi*, Jianshu Zhang*, Tianyang Han, Jipeng Zhang, Rui Pan, Tong Zhang
    ICLR 2025
  4. WebCoT: a web agent clicks the wrong link, rolls back, reflects and compares branches to find the right page, lifting WebVoyager success from 24.7% to 41.0% EMNLP
    Minda Hu, Tianqing Fang, Jianshu Zhang, Junyu Ma, Zhisong Zhang, Jingyan Zhou, Hongming Zhang, Haitao Mi, Dong Yu, Irwin King
    EMNLP 2025 (Findings); COLM 2025 Workshop AIA
  5. FairReason: sliding the training mix between reasoning and debiasing data trades accuracy for fairness, and a 1:4 mix trained with RL keeps 88% of reasoning while cutting stereotypes by 10% arXiv
    Zhenyu Pan, Yutong Zhang, Jianshu Zhang, Haoran Lu, Haozheng Luo, Yuwei Han, Philip S. Yu, Manling Li, Han Liu
    arXiv preprint arXiv:2507.23067
  6. CAN: the server alone mistakes a synthetic fox for a dog, so the client that knows foxes best steers the generator, and each client then replays what it forgets most ICML
  7. Bridge-Coder: a model that cannot jump straight from a task to D code crosses on a commented Python code-bridge, lifting pass@1 on R, D, Bash and Racket ACL
    Jipeng Zhang*, Jianshu Zhang*, Yuanzhe Li*, Renjie Pi, Rui Pan, Runtao Liu, Zheng Ziqiang, Tong Zhang
    ACL 2025 (Findings)

2024

  1. Image Textualization: vision experts fact-check an MLLM caption, strike the hallucinated window and add missing details, making descriptions longer and more accurate NeurIPS
    Renjie Pi*, Jianshu Zhang*, Jipeng Zhang, Rui Pan, Zhekai Chen, Tong Zhang
    NeurIPS 2024
  2. MLLM-Protector: a harm detector flags an MLLM's harmful reply and a detoxifier rewrites it into a refusal, cutting attack success from 79% to 6% while the frozen model keeps its accuracy EMNLP
    Renjie Pi*, Tianyang Han*, Jianshu Zhang*, Yueqi Xie, Rui Pan, Qing Lian, Hanze Dong, Jipeng Zhang, Tong Zhang
    EMNLP 2024 (Main)
  3. FuzzLLM: a slot machine mixes jailbreak templates, constraints and questions into test prompts, filling a vulnerability heatmap where combined attack classes crack GPT-4 far more often (5% to 38%) ICASSP
    Dongyu Yao*, Jianshu Zhang*, Ian G. Harris, Marcel Carlsson
    ICASSP 2024; Presented at ShmooCon 2024
  4. FIRST: keep the teacher's top-5 tokens, re-calibrate its over-confidence with temperature scaling, and distill a student that is confident when right and unsure when wrong EMNLP
    KaShun Shum*, Minrui Xu*, Jianshu Zhang*, Zixin Chen, Shizhe Diao, Hanze Dong, Jipeng Zhang, Muhammad Omer Raza
    EMNLP 2024 (Main)
  5. CORE: as new tasks arrive old memories fade, so the replay buffer gives the most space to the tasks forgotten fastest and replays them back up CogSci
    Jianshu Zhang*, Yankai Fu*, Ziheng Peng*, Dongyu Yao, Kun He
    CogSci 2024 (Oral)

Awards

  • McCormick School of Engineering Fellowship (~46K USD)
  • National Scholarship
  • Lei Jun Computer Breakthrough Award (50K RMB)
  • Outstanding Undergraduate
  • First-Class Scholarship (ranked 1st)
  • Merit Student

Education

  • Northwestern University, Ph.D. in Computer Science (2025-2030)
  • Wuhan University (2021-2025)
  • Shenzhen Middle School (2018-2021)