Skip to content
01

About Me

I am a Ph.D. student at Westlake University (Fall 2025), advised by Prof. Tailin Wu. I am deeply interested in the physics of intelligence. I do research to reduce my perplexity about the world we live in.

I mainly study the interplay between understanding—or representation learning—and generative models. I believe this interaction underlies many important phenomena in language and vision: why does changing the order of training data alter a language model's convergence speed, and why does the choice of representation space strongly affect the efficiency of visual generative modeling? I enjoy connecting these questions to statistical physics and information theory, and using scientific methods to uncover the principles behind them.

Previously, I studied astronomy at Nanjing University and the Australian National University, working with Prof. Yuan-Sen Ting.

02

Research Highlight

Submitted to NeurIPS 2026 Feature Information Dynamics in Diffusion Click to expand details

Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative—and, crucially, it need not hold after data are mapped into a representation space.

Feature information density

Let \(X\in\mathbb{R}^d\) be clean data (e.g., an image or latent), \(Y\) a feature of \(X\) (e.g., class, mask, or Canny), and \(X_\gamma=\sqrt{\gamma}X+N\) the Gaussian-corrupted data at signal-to-noise ratio \(\gamma\), where \(N\sim\mathcal{N}(0,I)\) is independent noise. We define the feature information density as

\[D_Y(\gamma):=\frac{\mathrm d}{\mathrm d\gamma}I(Y;X_\gamma).\]

Intuitively, \(D_Y(\gamma)\) distributes the total information about \(Y\) along the SNR axis: it measures how much additional feature information becomes accessible from an infinitesimal increase in SNR. A peak therefore identifies the noise level at which that feature is revealed most rapidly during denoising.

At any fixed SNR, a feature-conditional denoiser has access to \(Y\) in addition to \(X_\gamma\). Since it can always ignore this extra condition, its best achievable denoising loss \(m_Y(\gamma)\) cannot exceed the optimal unconditional loss \(m_\varnothing(\gamma)\). Using the I-MMSE identity, we show that feature information density is exactly half of this reduction in optimal denoising loss brought by feature conditioning:

\[D_Y(\gamma)=\frac{1}{2}\left[m_\varnothing(\gamma)-m_Y(\gamma)\right].\]

Its trajectory across noise levels describes how that feature's information is distributed over the generation process. Empirically, we find that class, mask, and Canny information exhibit markedly different dynamics across pixel, SDVAE, VAVAE, and RAE spaces. Among them, only RAE follows the class → mask → Canny order. We hypothesize that this ordered feature dynamics may explain why diffusion models converge fastest in the RAE space.

Feature information dynamics for class, mask, and Canny features in pixel space
Feature information dynamics in pixel space. Select the figure for the full-resolution PDF.
03

Selected Publications

NeurIPS 2026 · Under Review

Feature Information Dynamics in Diffusion

Jia-Shu Pan, Tao Zhang, Yufei Huang, Yanjun Sheng, Tailin Wu

A quantitative framework for locating hierarchical features along diffusion trajectories and relating their temporal organization to representation-dependent convergence.

ICLR 2026

VFScale: Intrinsic Reasoning through Verifier-Free Test-time Scalable Diffusion Model

Tao Zhang*, Jia-Shu Pan*, Ruiqi Feng, Tailin Wu

* Equal contribution

VFScale trains a diffusion model's own energy to serve as a verifier and combines it with hybrid Monte Carlo Tree Search, enabling verifier-free test-time scaling on Maze and Sudoku.

04

Hobbies

🏃 Running 🎞️ Anime
05

Contact

Interesting questions, disagreements, and half-formed ideas are always welcome. If our research interests overlap, I would be happy to talk.

Email: panjiashu@westlake.edu.cn

© 2026 Jia-Shu Pan · Built with Jekyll and al-folio · Website design inspired by Huanran Chen · Milky Way photographed by the author in the Tengger Desert on the night of August 12, 2026.