About
I am a PhD student at LPSM, Sorbonne University, and Google DeepMind, working on diffusion models and machine learning theory.
- Advisors
- Gérard Biau, Claire Boyer and Pierre Marion
- At Google DeepMind
- Quentin Berthet and Romuald Elie
Research focus
Generalization in diffusion models
Why diffusion models produce novel samples rather than copies of their training data, and how optimization shapes this.
Rethinking generalization · Optimal Stopping · Taking a Big Step
All 3 papersSampling and evaluation
When to stop the reverse process in latent diffusion models, and how to measure sample quality reliably.
All 2 papersGenerative models for science and language
Generative emulators for PDE simulations, and discrete-diffusion language models.
All 2 papersImplicit regularization in deep learning
How gradient descent biases deep residual networks toward neural ODEs, and how large learning rates regularize score matching.
Taking a Big Step · ResNets and neural ODEs
All 2 papersSelected publications
arXiv preprint · 2026
Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulation
Fine-tunes a deterministic physics foundation model into a fast, accurate generative PDE emulator.
- Problem. Autoregressive ML emulators of PDEs accumulate error over long horizons and miss the stochasticity of physical systems.
- Approach. Adapt a deterministic physics foundation model into a generative surrogate with two-stage inference, mean-prediction regularization and spatial gradient matching.
- Result. 42.9% average reduction in forecasting error over the Walrus fine-tuning baseline on The Well, better on 8 of 10 datasets.
NeurIPS 2026 position track
Understanding diffusion models requires rethinking (again) generalization
Generalization in diffusion models needs new theory: what is learned before memorization?
- Problem. In diffusion models memorization and generalization are incompatible, so the supervised-learning theory of generalization does not transfer.
- Approach. Survey the three families of explanations and run controlled CIFAR-10 sweeps over dataset size, model size, batch size and learning rate.
- Position. Why models do not memorize is largely settled by early stopping and the linear scaling of memorization time; what is learned before memorization is the open question.
ICML 2026 Oral at the PriGM workshop, EurIPS 2025
Optimal Stopping in Latent Diffusion Model
Why the last denoising steps of a latent diffusion model can hurt, and when to stop.
- Problem. The last steps of latent diffusion can degrade samples, which does not happen in pixel-space diffusion.
- Approach. A Gaussian model with linear autoencoders that links latent dimension, the constraints of score matching and the stopping time.
- Result. Low-dimensional latents call for earlier stopping, high-dimensional ones for later; stopping time is a key hyperparameter of latent diffusion.
COLT 2025
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization
Large learning rates implicitly regularize denoising score matching and prevent memorization.
- Problem. The exact minimizer of denoising score matching memorizes the training data, yet trained models memorize only mildly.
- Approach. Analyze gradient descent with large steps, which can only converge to minima whose sharpness is bounded by the inverse learning rate.
- Result. The memorizing solution is too sharp to be reached; large learning rates prevent memorization.
News
- New preprint: Kastor, an efficient fine-tuning strategy for generative emulation of PDE simulations, with colleagues at Google DeepMind.
- The DiffusionGemma technical report is out: an open-weight discrete-diffusion language model obtained by fine-tuning Gemma 4.
- Talk at MathSTIC (USPN) on Understanding diffusion models requires rethinking (again) generalization.
- Two new preprints: Understanding diffusion models requires rethinking (again) generalization (with Pierre Marion) and MIND, a sample-efficient alternative to FID.
- Optimal Stopping in Latent Diffusion Models accepted at ICML 2026.