About Me关于我
I am a Ph.D. student at Westlake University (Fall 2025), advised by Prof. Tailin Wu. I am deeply interested in the physics of intelligence. I do research to reduce my perplexity about the world we live in.
I mainly study the interplay between understanding—or representation learning—and generative models. I believe this interaction underlies many important phenomena in language and vision: why does changing the order of training data alter a language model's convergence speed, and why does the choice of representation space strongly affect the efficiency of visual generative modeling? I enjoy connecting these questions to statistical physics and information theory, and using scientific methods to uncover the principles behind them.
Previously, I studied astronomy at Nanjing University and the Australian National University, working with Prof. Yuan-Sen Ting.
我于 2025 年秋季进入西湖大学攻读博士学位,导师是吴泰霖老师。我对智能的物理学有浓厚兴趣。我做研究,是为了减少自己对我们所生活的世界的困惑。
我目前主要关注理解(或者说表示学习)与生成模型之间的相互作用。我相信这种相互作用支配了语言和视觉建模中众多有趣且重要的现象:为什么语言模型看到训练数据的顺序不同,会产生不同的收敛速度?为什么表示空间的选择会显著改变视觉生成模型的收敛效率?我享受这些问题与统计物理、信息论等领域的联系,并希望用科学方法理解现象背后的规律。
此前,我在南京大学和澳大利亚国立大学学习天文,并与Yuan-Sen Ting 教授合作。
Research Highlight研究工作
Submitted to NeurIPS 2026 Feature Information Dynamics in Diffusion 投稿于 NeurIPS 2026 扩散模型中的特征信息动力学
Diffusion models generate data through a continuum of denoising problems, and are widely observed to reveal coarse structure before fine detail. Yet, this intuition is mostly empirical and qualitative—and, crucially, it need not hold after data are mapped into a representation space.
Feature information density
Let \(X\in\mathbb{R}^d\) be clean data (e.g., an image or latent), \(Y\) a feature of \(X\) (e.g., class, mask, or Canny), and \(X_\gamma=\sqrt{\gamma}X+N\) the Gaussian-corrupted data at signal-to-noise ratio \(\gamma\), where \(N\sim\mathcal{N}(0,I)\) is independent noise. We define the feature information density as
Intuitively, \(D_Y(\gamma)\) distributes the total information about \(Y\) along the SNR axis: it measures how much additional feature information becomes accessible from an infinitesimal increase in SNR. A peak therefore identifies the noise level at which that feature is revealed most rapidly during denoising.
At any fixed SNR, a feature-conditional denoiser has access to \(Y\) in addition to \(X_\gamma\). Since it can always ignore this extra condition, its best achievable denoising loss \(m_Y(\gamma)\) cannot exceed the optimal unconditional loss \(m_\varnothing(\gamma)\). Using the I-MMSE identity, we show that feature information density is exactly half of this reduction in optimal denoising loss brought by feature conditioning:
Its trajectory across noise levels describes how that feature's information is distributed over the generation process. Empirically, we find that class, mask, and Canny information exhibit markedly different dynamics across pixel, SDVAE, VAVAE, and RAE spaces. Among them, only RAE follows the class → mask → Canny order. We hypothesize that this ordered feature dynamics may explain why diffusion models converge fastest in the RAE space.
扩散模型通过一系列连续的去噪问题生成数据,人们普遍观察到它会先呈现粗粒度结构,再补充细节。然而,这一直主要是一种经验性的定性直觉;更关键的是,当数据被映射到表示空间后,这一 coarse-to-fine 直觉未必仍然成立。
特征信息密度
设 \(X\in\mathbb{R}^d\) 是干净数据(例如图像或 latent),\(Y\) 是 \(X\) 的某个特征(例如类别、mask 或 Canny),\(X_\gamma=\sqrt{\gamma}X+N\) 是信噪比 \(\gamma\) 下经过高斯扰动的数据,其中 \(N\sim\mathcal{N}(0,I)\) 是独立噪声。我们将特征信息密度定义为
直观上,\(D_Y(\gamma)\) 将关于 \(Y\) 的总信息量分布到信噪比轴上:它衡量信噪比增加无穷小量时,我们能从数据中多获得多少关于该特征的信息。因此,曲线的峰值对应这一特征在去噪过程中显现得最快的噪声水平。
在任意固定的信噪比上,特征条件去噪器除了 \(X_\gamma\) 之外还能使用 \(Y\)。由于它总可以选择忽略这一额外条件,其能够达到的最优去噪损失 \(m_Y(\gamma)\) 不会高于无条件去噪器的最优损失 \(m_\varnothing(\gamma)\)。我们利用 I-MMSE 关系证明,特征信息密度恰好等于这一特征条件带来的最优去噪损失下降的一半:
它随噪声水平变化的轨迹刻画了该特征信息在生成过程中的分布。经验上,我们发现类别、mask 和 Canny 信息在像素、SDVAE、VAVAE 与 RAE 空间中呈现显著不同的动力学。其中,只有 RAE 上的特征信息动力学服从类别 → mask → Canny 的顺序。我们猜测,这种有序的特征动力学可能正是扩散模型在 RAE 空间中收敛最快的原因。
Selected Publications代表论文
NeurIPS 2026 · Under Review
Feature Information Dynamics in Diffusion
A quantitative framework for locating hierarchical features along diffusion trajectories and relating their temporal organization to representation-dependent convergence.
定量定位层次特征在扩散轨迹中的生成时刻,并研究这种时间组织方式与不同表示空间收敛速度之间的关系。
ICLR 2026
VFScale: Intrinsic Reasoning through Verifier-Free Test-time Scalable Diffusion Model
* Equal contribution* 共同第一作者
VFScale trains a diffusion model's own energy to serve as a verifier and combines it with hybrid Monte Carlo Tree Search, enabling verifier-free test-time scaling on Maze and Sudoku.
VFScale 将扩散模型自身的能量训练为验证器,并结合混合蒙特卡洛树搜索,在迷宫与数独任务上实现无需外部验证器的测试时扩展。
Hobbies爱好
Contact联系我
Interesting questions, disagreements, and half-formed ideas are always welcome. If our research interests overlap, I would be happy to talk.
Email: panjiashu@westlake.edu.cn
有趣的问题、不同的意见和尚未成形的想法都很受欢迎。如果我们的研究兴趣有所交集,欢迎随时交流。
© 2026 Jia-Shu Pan · Built with Jekyll and al-folio · Website design inspired by Huanran Chen · Milky Way photographed by the author in the Tengger Desert on the night of August 12, 2026. © 2026 潘嘉书 · 使用 Jekyll 与 al-folio 构建 · 网站设计参考了陈焕然的个人主页 · 星空照片由本人于 2026 年 8 月 12 日晚摄于腾格里沙漠。